{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"---\n# [House Prices - Advanced Regression Techniques][1]\n\n- Goal: It is your job to predict the sales price for each house. For each Id in the test set, you must predict the value of the SalePrice variable. \n\n---\n#### **The aim of this notebook is to**\n- **1. Conduct Exploratory Data Analysis (EDA) and Feature Engineering.**\n- **2. Use PyCaret for the entire ML pipeline.**\n- **3. Create and train a Deep Learning model with TensorFlow.**\n\n---\n**References:** Thanks to previous great codes and notebooks.\n\n- [PyCaret Tutoricals][2]\n - [Regression Tutorial - Level Beginner][3]\n - [Regression Tutorial  - Level Intermediate][4]\n \n**My Previous Notebooks:**\n- [SpaceshipTitanic: EDA + TabTransformer[TensorFlow]][5]\n\n---\n### **If you find this notebook useful, please do give me an upvote. It helps me keep up my motivation.**\n#### **Also, I would appreciate it if you find any mistakes and help me correct them.**\n\n---\n[1]: https://www.kaggle.com/competitions/house-prices-advanced-regression-techniques\n[2]: https://pycaret.gitbook.io/docs/get-started/tutorials\n[3]: https://github.com/pycaret/pycaret/blob/master/tutorials/Regression%20Tutorial%20Level%20Beginner%20-%20REG101.ipynb\n[4]: https://github.com/pycaret/pycaret/blob/master/tutorials/Regression%20Tutorial%20Level%20Intermediate%20-%20REG102.ipynb\n[5]: https://www.kaggle.com/code/masatomurakawamm/spaceshiptitanic-eda-tabtransformer-tensorflow","metadata":{}},{"cell_type":"markdown","source":"<h1 style=\"background:#05445E; border:0; border-radius: 12px; color:#D3D3D3\"><center>0. TABLE OF CONTENTS</center></h1>\n\n<ul class=\"list-group\" style=\"list-style-type:none;\">\n    <li><a href=\"#1\" class=\"list-group-item list-group-item-action\">1. Settings</a></li>\n    <li><a href=\"#2\" class=\"list-group-item list-group-item-action\">2. Data Loading</a></li>\n    <li><a href=\"#3\" class=\"list-group-item list-group-item-action\">3. EDA and Feature Engineering</a>\n        <ul class=\"list-group\" style=\"list-style-type:none;\">\n            <li><a href=\"#3.1\" class=\"list-group-item list-group-item-action\">3.1 AutoEDA with Sweetviz</a></li>\n            <li><a href=\"#3.2\" class=\"list-group-item list-group-item-action\">3.2 Feature Selection</a></li>\n            <li><a href=\"#3.3\" class=\"list-group-item list-group-item-action\">3.3 Target Distribution</a></li>\n            <li><a href=\"#3.4\" class=\"list-group-item list-group-item-action\">3.4 Numerical Features</a></li>\n            <li><a href=\"#3.5\" class=\"list-group-item list-group-item-action\">3.5 Categorical Feature</a></li>\n        </ul>\n    </li>\n    <li><a href=\"#4\" class=\"list-group-item list-group-item-action\">4. PyCaret</a>\n        <ul class=\"list-group\" style=\"list-style-type:none;\">\n            <li><a href=\"#4.1\" class=\"list-group-item list-group-item-action\">4.1 Setting up Environment</a></li>\n            <li><a href=\"#4.2\" class=\"list-group-item list-group-item-action\">4.2 Create Model</a></li>\n            <li><a href=\"#4.3\" class=\"list-group-item list-group-item-action\">4.3 Tune Model</a></li>\n            <li><a href=\"#4.4\" class=\"list-group-item list-group-item-action\">4.4 Plot Model</a></li>\n            <li><a href=\"#4.5\" class=\"list-group-item list-group-item-action\">4.5 Validate Model on Hold-out Sample</a></li>\n            <li><a href=\"#4.6\" class=\"list-group-item list-group-item-action\">4.6 Finalize Model and Inference</a></li>\n            <li><a href=\"#4.7\" class=\"list-group-item list-group-item-action\">4.7 Ensemble Models</a></li>\n        </ul>\n    </li>\n    <li><a href=\"#5\" class=\"list-group-item list-group-item-action\">5. Deep Learning</a>\n        <ul class=\"list-group\" style=\"list-style-type:none;\">\n            <li><a href=\"#5.1\" class=\"list-group-item list-group-item-action\">5.1 Creating Dataset</a></li>\n            <li><a href=\"#5.2\" class=\"list-group-item list-group-item-action\">5.2 Preprocessing Model</a></li>\n            <li><a href=\"#5.3\" class=\"list-group-item list-group-item-action\">5.3 Training Model</a></li>\n            <li><a href=\"#5.4\" class=\"list-group-item list-group-item-action\">5.4 Model Training</a></li>\n            <li><a href=\"#5.5\" class=\"list-group-item list-group-item-action\">5.5 Inference</a></li>\n        </ul>\n    </li>\n</ul>","metadata":{}},{"cell_type":"markdown","source":"<a id =\"1\"></a><h1 style=\"background:#05445E; border:0; border-radius: 12px; color:#D3D3D3\"><center>1. Settings</center></h1>","metadata":{}},{"cell_type":"code","source":"## Import dependencies \nimport numpy as np\nimport pandas as pd\nimport scipy as sp\nimport matplotlib.pyplot as plt \n%matplotlib inline\n\nimport seaborn as sns\nimport plotly.express as px\nimport plotly.graph_objects as go\nfrom plotly.subplots import make_subplots\n\nimport os\nimport pathlib\nimport gc\nimport sys\nimport re\nimport math \nimport random\nimport time \nimport datetime as dt\nfrom tqdm import tqdm \n\nprint('Import done!')","metadata":{"execution":{"iopub.status.busy":"2022-07-25T12:07:13.268213Z","iopub.execute_input":"2022-07-25T12:07:13.268600Z","iopub.status.idle":"2022-07-25T12:07:16.139481Z","shell.execute_reply.started":"2022-07-25T12:07:13.268483Z","shell.execute_reply":"2022-07-25T12:07:16.138540Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## For reproducible results    \ndef seed_all(s):\n    random.seed(s)\n    np.random.seed(s)\n    os.environ['PYTHONHASHSEED'] = str(s) \n    print('Seeds setted!')\n    \nglobal_seed = 42\nseed_all(global_seed)","metadata":{"execution":{"iopub.status.busy":"2022-07-25T12:07:16.141283Z","iopub.execute_input":"2022-07-25T12:07:16.142131Z","iopub.status.idle":"2022-07-25T12:07:16.149189Z","shell.execute_reply.started":"2022-07-25T12:07:16.142082Z","shell.execute_reply":"2022-07-25T12:07:16.148128Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Parameters\ndata_config = {'train_csv_path': '../input/house-prices-advanced-regression-techniques/train.csv',\n               'test_csv_path': '../input/house-prices-advanced-regression-techniques/test.csv',\n               'sample_submission_path': '../input/house-prices-advanced-regression-techniques/sample_submission.csv',\n              }\n\nexp_config = {'gpu': True,\n              'n_splits': 5,\n              'batch_size': 128,\n              'learning_rate': 1e-3,\n              'train_epochs': 100,\n              'checkpoint_filepath': './tmp/model/exp.ckpt',\n             }\n\nmodel_config = {'emb_dim': 3,\n                'model_units': [512, 128],\n                'dropout_rates': [0.2, 0.2,],\n               }\n\nprint('Parameters setted!')","metadata":{"execution":{"iopub.status.busy":"2022-07-25T12:07:16.150766Z","iopub.execute_input":"2022-07-25T12:07:16.151079Z","iopub.status.idle":"2022-07-25T12:07:16.164676Z","shell.execute_reply.started":"2022-07-25T12:07:16.151037Z","shell.execute_reply":"2022-07-25T12:07:16.163422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =\"2\"></a><h1 style=\"background:#05445E; border:0; border-radius: 12px; color:#D3D3D3\"><center>2. Data Loading</center></h1>","metadata":{}},{"cell_type":"markdown","source":"---\n### [File and Data Field Descriptions](https://www.kaggle.com/competitions/house-prices-advanced-regression-techniques/data)\n\n- **train.csv** - the training set\n- **test.csv** - the test set\n- **data_description.txt** - full description of each column, originally prepared by Dean De Cock but lightly edited to match the column names used here\n- **sample_submission.csv** - a benchmark submission from a linear regression on year and month of sale, lot square footage, and number of bedrooms.\n\n\n---\n### [Submission & Evaluation](https://www.kaggle.com/competitions/house-prices-advanced-regression-techniques/overview/evaluation)\n\n- Submissions are evaluated on Root-Mean-Squared-Error (RMSE) between the logarithm of the predicted value and the logarithm of the observed sales price. (Taking logs means that errors in predicting expensive houses and cheap houses will affect the result equally.)\n\n---","metadata":{}},{"cell_type":"code","source":"## Data Loading\ntrain_df = pd.read_csv(data_config['train_csv_path'])\ntest_df = pd.read_csv(data_config['test_csv_path'])\nsubmission_df = pd.read_csv(data_config['sample_submission_path'])\n\nprint(f'train_length: {len(train_df)}')\nprint(f'test_lenght: {len(test_df)}')\nprint(f'submission_length: {len(submission_df)}')","metadata":{"execution":{"iopub.status.busy":"2022-07-25T12:07:16.167151Z","iopub.execute_input":"2022-07-25T12:07:16.167682Z","iopub.status.idle":"2022-07-25T12:07:16.299522Z","shell.execute_reply.started":"2022-07-25T12:07:16.167635Z","shell.execute_reply":"2022-07-25T12:07:16.298258Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Null Value Check\nprint('train_df.info()'); print(train_df.info(), '\\n')\nprint('test_df.info()'); print(test_df.info(), '\\n')\n\n## train_df Check\ntrain_df.head()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:07:16.301871Z","iopub.execute_input":"2022-07-25T12:07:16.303009Z","iopub.status.idle":"2022-07-25T12:07:16.443306Z","shell.execute_reply.started":"2022-07-25T12:07:16.302819Z","shell.execute_reply":"2022-07-25T12:07:16.441527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =\"3\"></a><h1 style=\"background:#05445E; border:0; border-radius: 12px; color:#D3D3D3\"><center>3. Exploratory Data Analysis</center></h1>","metadata":{}},{"cell_type":"markdown","source":"<a id =\"3.1\"></a><h2 style=\"background:#75E6DA; border:0; border-radius: 12px; color:black\"><center>3.1 AutoEDA with Sweetviz</center></h2>","metadata":{}},{"cell_type":"code","source":"## Import dependencies\n!pip install -U -q sweetviz \nimport sweetviz","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:07:16.446234Z","iopub.execute_input":"2022-07-25T12:07:16.447337Z","iopub.status.idle":"2022-07-25T12:07:31.652429Z","shell.execute_reply.started":"2022-07-25T12:07:16.447238Z","shell.execute_reply":"2022-07-25T12:07:31.651474Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"my_report_1 = sweetviz.analyze(train_df)\nmy_report_1.show_notebook()","metadata":{"execution":{"iopub.status.busy":"2022-07-25T12:07:31.654084Z","iopub.execute_input":"2022-07-25T12:07:31.654392Z","iopub.status.idle":"2022-07-25T12:08:21.352184Z","shell.execute_reply.started":"2022-07-25T12:07:31.654348Z","shell.execute_reply":"2022-07-25T12:08:21.351186Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"my_report_2 = sweetviz.compare([train_df, \"Train\"], [test_df, \"Test\"], \"SalePrice\")\nmy_report_2.show_notebook()","metadata":{"execution":{"iopub.status.busy":"2022-07-25T12:08:21.353580Z","iopub.execute_input":"2022-07-25T12:08:21.353961Z","iopub.status.idle":"2022-07-25T12:09:54.675156Z","shell.execute_reply.started":"2022-07-25T12:08:21.353915Z","shell.execute_reply":"2022-07-25T12:09:54.673884Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =\"3.1\"></a><h2 style=\"background:#75E6DA; border:0; border-radius: 12px; color:black\"><center>3.1 Feature Selection</center></h2>","metadata":{}},{"cell_type":"code","source":"## Drop the columns which contain null values more than 500.\ndef feature_selection(dataframe):\n    df = dataframe.copy()\n    for column in df.columns:\n        n_null = train_df[column].isnull().sum()\n        if n_null > 500:\n            df = df.drop([column], axis=1)\n    return df\n\n## Feature selection on train data\ntrain = feature_selection(train_df)\nprint(len(train_df.columns), len(train.columns))","metadata":{"execution":{"iopub.status.busy":"2022-07-25T12:09:54.676848Z","iopub.execute_input":"2022-07-25T12:09:54.677527Z","iopub.status.idle":"2022-07-25T12:09:54.728478Z","shell.execute_reply.started":"2022-07-25T12:09:54.677478Z","shell.execute_reply":"2022-07-25T12:09:54.727370Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Feature selection on test data\nfeature_list = list(train.columns)\nfeature_list.remove('SalePrice')\ntest = test_df[feature_list]","metadata":{"execution":{"iopub.status.busy":"2022-07-25T12:09:54.732444Z","iopub.execute_input":"2022-07-25T12:09:54.733330Z","iopub.status.idle":"2022-07-25T12:09:54.739993Z","shell.execute_reply.started":"2022-07-25T12:09:54.733287Z","shell.execute_reply":"2022-07-25T12:09:54.738942Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Select numerical and categorical features\ndef get_numerical_categorical(df, feature_list, n_unique=False):\n    numerical_features = []\n    categorical_features = []\n    for column in feature_list:\n        if n_unique:\n            if df[column].nunique() > n_unique:\n                numerical_features.append(column)\n            else:\n                categorical_features.append(column)\n        else:\n            if df[column].dtypes == 'object': \n                categorical_features.append(column)\n            else:\n                numerical_features.append(column)\n    return numerical_features, categorical_features\n\ntarget = 'SalePrice'\n\n## Features which has more than 30 unique values as numerical\nnumerical_features, categorical_features = get_numerical_categorical(train,\n                                                                     feature_list,\n                                                                     n_unique=30)\nnumerical_features.remove('Id')\nprint(len(numerical_features), len(categorical_features))","metadata":{"execution":{"iopub.status.busy":"2022-07-25T12:09:54.741354Z","iopub.execute_input":"2022-07-25T12:09:54.741631Z","iopub.status.idle":"2022-07-25T12:09:54.774284Z","shell.execute_reply.started":"2022-07-25T12:09:54.741592Z","shell.execute_reply":"2022-07-25T12:09:54.773348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Numerical features' dtype check \nfor n in range(len(numerical_features)):\n    print(numerical_features[n])\n    print(train[numerical_features[n]].dtypes)\n    print()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:09:54.775702Z","iopub.execute_input":"2022-07-25T12:09:54.775934Z","iopub.status.idle":"2022-07-25T12:09:54.787477Z","shell.execute_reply.started":"2022-07-25T12:09:54.775905Z","shell.execute_reply":"2022-07-25T12:09:54.786777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Categoical features' unique values check\nfor n in range(len(categorical_features)):\n    print(categorical_features[n])\n    print(train[categorical_features[n]].unique())\n    print()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:09:54.788850Z","iopub.execute_input":"2022-07-25T12:09:54.789288Z","iopub.status.idle":"2022-07-25T12:09:54.830936Z","shell.execute_reply.started":"2022-07-25T12:09:54.789246Z","shell.execute_reply":"2022-07-25T12:09:54.829906Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =\"3.2\"></a><h2 style=\"background:#75E6DA; border:0; border-radius: 12px; color:black\"><center>3.2 Target Distribution</center></h2>","metadata":{}},{"cell_type":"code","source":"sns.histplot(x=target, data=train, kde=True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-25T12:09:54.832141Z","iopub.execute_input":"2022-07-25T12:09:54.832383Z","iopub.status.idle":"2022-07-25T12:09:55.188540Z","shell.execute_reply.started":"2022-07-25T12:09:54.832352Z","shell.execute_reply":"2022-07-25T12:09:55.187736Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### [Box-Cox Transformation](https://qiita.com/dyamaguc/items/b468ae66f9ce6ee89724) on 'SalePrice'","metadata":{}},{"cell_type":"code","source":"fig = plt.figure(figsize=(10, 8))\n\nlist_lambda = [-2, -1, -0.5, 0, 0.5, 1, 2]\nfor i, lmbda in enumerate(list_lambda):\n    boxcox = sp.stats.boxcox(train[target], lmbda=lmbda)\n    ax = fig.add_subplot(4, 2, i+1)\n    sns.histplot(data=boxcox, kde=True, ax=ax)\n    plt.title('lambda='+str(list_lambda[i]))\n    plt.xlabel('SalePrice')\n    \nauto_boxcox, best_lambda = sp.stats.boxcox(train[target], lmbda=None)\nax = fig.add_subplot(4, 2, 8)\nsns.histplot(data=auto_boxcox, kde=True, ax=ax)\nplt.title('lambda=' + str(round(best_lambda, 2)))\nplt.xlabel('SalePrice')\n    \nfig.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2022-07-25T12:09:55.189990Z","iopub.execute_input":"2022-07-25T12:09:55.190217Z","iopub.status.idle":"2022-07-25T12:09:57.675530Z","shell.execute_reply.started":"2022-07-25T12:09:55.190189Z","shell.execute_reply":"2022-07-25T12:09:57.674813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Box-Cox Summary\nfig = plt.figure(figsize=(12, 5))\n\n## Box-Cox Transformation\nauto_boxcox, best_lambda = sp.stats.boxcox(train[target], lmbda=None)\nax = fig.add_subplot(1, 2, 1)\nsns.histplot(data=auto_boxcox, kde=True, ax=ax)\nplt.title('Transformed')\nplt.xlabel('SalePrice')\n\n## Reverse Transformation\nx = sp.special.inv_boxcox(auto_boxcox, best_lambda)\nax = fig.add_subplot(1, 2, 2)\nsns.histplot(x=x, data=train, kde=True, ax=ax)\nplt.title('Reverse Transformed')\nplt.xlabel('SalePrice')\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-25T12:09:57.676665Z","iopub.execute_input":"2022-07-25T12:09:57.677031Z","iopub.status.idle":"2022-07-25T12:09:58.260219Z","shell.execute_reply.started":"2022-07-25T12:09:57.677002Z","shell.execute_reply":"2022-07-25T12:09:58.259551Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =\"3.3\"></a><h2 style=\"background:#75E6DA; border:0; border-radius: 12px; color:black\"><center>3.3 Numerical Features</center></h2>","metadata":{}},{"cell_type":"code","source":"## Distributions of numerical features\nfig = plt.figure(figsize=(10, 18))\nfor i, nf in enumerate(numerical_features):\n    ax = fig.add_subplot(6, 3, i+1)\n    sns.histplot(train[nf], kde=True, ax=ax)\n    plt.title(nf)\n    plt.xlabel(None)\nfig.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2022-07-25T12:09:58.261348Z","iopub.execute_input":"2022-07-25T12:09:58.261724Z","iopub.status.idle":"2022-07-25T12:10:05.663147Z","shell.execute_reply.started":"2022-07-25T12:09:58.261690Z","shell.execute_reply":"2022-07-25T12:10:05.662196Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Heatmap of correlation matrix\nnumerical_columns = numerical_features + ['SalePrice']\ntrain_numerical = train[numerical_columns]\n\nfig = px.imshow(train_numerical.corr(),\n                color_continuous_scale='RdBu_r',\n                color_continuous_midpoint=0, \n                aspect='auto')\nfig.update_layout(height=600, \n                  width=600,\n                  title = \"Heatmap\",                  \n                  showlegend=False)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-25T12:10:05.664764Z","iopub.execute_input":"2022-07-25T12:10:05.665182Z","iopub.status.idle":"2022-07-25T12:10:06.807338Z","shell.execute_reply.started":"2022-07-25T12:10:05.665130Z","shell.execute_reply":"2022-07-25T12:10:06.806523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =\"3.4\"></a><h2 style=\"background:#75E6DA; border:0; border-radius: 12px; color:black\"><center>3.4 Categorical Features</center></h2>","metadata":{}},{"cell_type":"code","source":"## Distributions of categorical features\nfig = plt.figure(figsize=(10, 50))\nfor i, cf in enumerate(categorical_features):\n    ax = fig.add_subplot(19, 3, i+1)\n    sns.histplot(train[cf], kde=False, ax=ax)\n    plt.title(cf)\n    plt.xlabel(None)\nfig.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2022-07-25T12:10:06.809123Z","iopub.execute_input":"2022-07-25T12:10:06.809701Z","iopub.status.idle":"2022-07-25T12:10:16.613202Z","shell.execute_reply.started":"2022-07-25T12:10:06.809655Z","shell.execute_reply":"2022-07-25T12:10:16.612542Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =\"4\"></a><h1 style=\"background:#05445E; border:0; border-radius: 12px; color:#D3D3D3\"><center>4. PyCaret</center></h1>","metadata":{}},{"cell_type":"markdown","source":"#### [PyCaret](https://pycaret.org/) is an open-source, low-code machine learning library in Python that automates machine learning workflows.","metadata":{}},{"cell_type":"code","source":"## Installing and importing dependencies\n!pip install -U -q pycaret --ignore-installed llvmlite\n#!pip install -U -q --pre pycaret\n!pip install -U -q numba==0.53 --ignore-installed llvmlite\n!pip install -U -q Pillow==9.1.0\n\nfrom pycaret.regression import *\nprint('Import done!')","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:10:16.614495Z","iopub.execute_input":"2022-07-25T12:10:16.614905Z","iopub.status.idle":"2022-07-25T12:11:11.632506Z","shell.execute_reply.started":"2022-07-25T12:10:16.614869Z","shell.execute_reply":"2022-07-25T12:11:11.631382Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =\"4.1\"></a><h2 style=\"background:#75E6DA; border:0; border-radius: 12px; color:black\"><center>4.1 Setting up Environment</center></h2>","metadata":{}},{"cell_type":"markdown","source":"- The `setup()` function initializes the environment in pycaret and creates the transformation pipeline to prepare the data for modeling and deployment.\n- When `setup()` is executed, PyCaret will automatically infer the data types for all features, and displays a table containing the features and their inferred data types. If all of the data types are correctly identified `enter` can be pressed to continue or `quit` can be typed to end the expriment. When `silent` parameter equals `True`, you can omit the manual check.\n- Also, you can pass directly `numeric_features` or `categorical_features` parameters.","metadata":{}},{"cell_type":"code","source":"silent = True\n\nexp_reg123 = setup(data=train, \n                   target='SalePrice',\n                   train_size=0.8,\n                   #numeric_features=numerical_features,\n                   numeric_imputation='mean',\n                   #categorical_features=categorical_features,\n                   categorical_imputation='constant',\n                   handle_unknown_categorical=True,\n                   ordinal_features=None,\n                   date_features=None,\n                   ignore_features=['Id'],\n                   normalize=True,\n                   transformation=False,\n                   transform_target=True,\n                   transform_target_method='box-cox',\n                   combine_rare_levels=True,\n                   rare_level_threshold=0.05,\n                   remove_multicollinearity=True,\n                   multicollinearity_threshold=0.95, \n                   bin_numeric_features=None,\n                   log_experiment=False,\n                   experiment_name='house_prices_exp1',\n                   session_id=123,\n                   use_gpu=exp_config['gpu'],\n                   silent=silent) ","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:11:11.634305Z","iopub.execute_input":"2022-07-25T12:11:11.635272Z","iopub.status.idle":"2022-07-25T12:11:11.977736Z","shell.execute_reply.started":"2022-07-25T12:11:11.635228Z","shell.execute_reply":"2022-07-25T12:11:11.974028Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =\"4.2\"></a><h2 style=\"background:#75E6DA; border:0; border-radius: 12px; color:black\"><center>4.2 Create Model</center></h2>","metadata":{}},{"cell_type":"markdown","source":"- The `compare_models()` function conducts training and evaluation over 15 models using cross validation.\n- The default `fold` parameter value is 10. To seve time, I setted the parameter `fold=5` for saving time.\n- By default, `compare_models()` return the best performing model, but can be used to return a list of top N models by using `n_select` parameter.\n- `exclude` parameter is used to block certain models.\n- To create each models, we can use `create_model()` function.","metadata":{}},{"cell_type":"code","source":"top3 = compare_models(fold=5,\n                      n_select=3,\n                      round=2,\n                      exclude=['ransac'])","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:11:11.978865Z","iopub.status.idle":"2022-07-25T12:11:11.979224Z","shell.execute_reply.started":"2022-07-25T12:11:11.979043Z","shell.execute_reply":"2022-07-25T12:11:11.979062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Select the best model\nbest_model = top3[0]\n\nfor i in range(len(top3)):\n    print(f'Top {i+1} Model: \\n{top3[i]}\\n\\n' )","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:11:11.980851Z","iopub.status.idle":"2022-07-25T12:11:11.981558Z","shell.execute_reply.started":"2022-07-25T12:11:11.981359Z","shell.execute_reply":"2022-07-25T12:11:11.981380Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Create the specific model.\ncatboost = create_model('catboost', fold=5, round=2)\net = create_model('et', fold=5, round=2, verbose=False)\nrf = create_model('rf', fold=5, round=2, verbose=False)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:11:11.982607Z","iopub.status.idle":"2022-07-25T12:11:11.983126Z","shell.execute_reply.started":"2022-07-25T12:11:11.982946Z","shell.execute_reply":"2022-07-25T12:11:11.982966Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"models = [catboost, et, rf]\n\nfor model in models:\n    print(model, '\\n')","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:11:11.984187Z","iopub.status.idle":"2022-07-25T12:11:11.984743Z","shell.execute_reply.started":"2022-07-25T12:11:11.984542Z","shell.execute_reply":"2022-07-25T12:11:11.984564Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =\"4.3\"></a><h2 style=\"background:#75E6DA; border:0; border-radius: 12px; color:black\"><center>4.3 Tune Model</center></h2>","metadata":{}},{"cell_type":"markdown","source":"-  The `tune_model()` function automatically tunes the hyperparameters of a model using Random Grid Search on a pre-defined search space.\n- We can change the metric for optimization by `optimize` parameter (default, Accuracy).\n- `n_iter` is the number of iterations within a random grid search (default value is 10). Increasing the value may improve the performance but will also increase the training time.","metadata":{}},{"cell_type":"code","source":"## It will take some time.\n#tuned_best = tune_model(best_model, fold=5, n_iter=30)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:11:11.985798Z","iopub.status.idle":"2022-07-25T12:11:11.986335Z","shell.execute_reply.started":"2022-07-25T12:11:11.986134Z","shell.execute_reply":"2022-07-25T12:11:11.986158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =\"4.4\"></a><h2 style=\"background:#75E6DA; border:0; border-radius: 12px; color:black\"><center>4.4 Plot Model</center></h2>","metadata":{}},{"cell_type":"markdown","source":"- The `plot_model()` function can be used to analyze the performance across different aspects such as Residuals Plot, Prediction Error, Feature Importance etc. ","metadata":{}},{"cell_type":"code","source":"plot_model(catboost, plot='parameter')","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:11:11.987282Z","iopub.status.idle":"2022-07-25T12:11:11.987642Z","shell.execute_reply.started":"2022-07-25T12:11:11.987436Z","shell.execute_reply":"2022-07-25T12:11:11.987476Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_model(catboost, plot='residuals') ## Default","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:11:11.989236Z","iopub.status.idle":"2022-07-25T12:11:11.989921Z","shell.execute_reply.started":"2022-07-25T12:11:11.989744Z","shell.execute_reply":"2022-07-25T12:11:11.989764Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_model(catboost, plot = 'error')","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:11:11.990961Z","iopub.status.idle":"2022-07-25T12:11:11.991511Z","shell.execute_reply.started":"2022-07-25T12:11:11.991296Z","shell.execute_reply":"2022-07-25T12:11:11.991320Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_model(catboost, plot='feature')","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:11:11.992556Z","iopub.status.idle":"2022-07-25T12:11:11.993101Z","shell.execute_reply.started":"2022-07-25T12:11:11.992911Z","shell.execute_reply":"2022-07-25T12:11:11.992934Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- We can also use `evaluate_model()` function for further model analysis.","metadata":{}},{"cell_type":"code","source":"#evaluate_model(catboost)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:11:11.994092Z","iopub.status.idle":"2022-07-25T12:11:11.994662Z","shell.execute_reply.started":"2022-07-25T12:11:11.994430Z","shell.execute_reply":"2022-07-25T12:11:11.994453Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =\"4.5\"></a><h2 style=\"background:#75E6DA; border:0; border-radius: 12px; color:black\"><center>4.5 Validate Model on Hold-out Sample</center></h2>","metadata":{}},{"cell_type":"markdown","source":"-  Before finalizing the model, we can conduct the final check by predicting the hold-out set and reviewing the evaluation metrics (All of the evaluation metrics we have seen above are cross-validated results based on training set (80%) only).\n- Remaining 20% of data (hold-out samples) is used for `predict_model()` function.","metadata":{}},{"cell_type":"code","source":"predict_model(catboost)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:11:11.995690Z","iopub.status.idle":"2022-07-25T12:11:11.996215Z","shell.execute_reply.started":"2022-07-25T12:11:11.996027Z","shell.execute_reply":"2022-07-25T12:11:11.996050Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =\"4.6\"></a><h2 style=\"background:#75E6DA; border:0; border-radius: 12px; color:black\"><center>4.6 Finalize Model and Inference</center></h2>","metadata":{}},{"cell_type":"markdown","source":"- `finalize_model()` function fits the model onto the complete dataset including the hold-out sample. ","metadata":{}},{"cell_type":"code","source":"final_model = finalize_model(catboost)\nprint(final_model)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:11:11.997211Z","iopub.status.idle":"2022-07-25T12:11:11.997753Z","shell.execute_reply.started":"2022-07-25T12:11:11.997562Z","shell.execute_reply":"2022-07-25T12:11:11.997586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Inference\ntest_predictions = predict_model(final_model, data=test)\nsubmission_df['SalePrice'] = test_predictions['Label']\nsubmission_df.to_csv('submission_pycaret.csv', index=False)\ntest_predictions","metadata":{"_kg_hide-input":false,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:11:11.998749Z","iopub.status.idle":"2022-07-25T12:11:11.999264Z","shell.execute_reply.started":"2022-07-25T12:11:11.999073Z","shell.execute_reply":"2022-07-25T12:11:11.999095Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =\"4.7\"></a><h2 style=\"background:#75E6DA; border:0; border-radius: 12px; color:black\"><center>4.7 Ensemble Models</center></h2>","metadata":{}},{"cell_type":"markdown","source":"- `blend_model()` function creates multiple models and then averages the individual predictions to form a final prediction. ","metadata":{}},{"cell_type":"code","source":"## Blend individual models\nblender = blend_models(estimator_list=models, fold=5)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:11:12.000240Z","iopub.status.idle":"2022-07-25T12:11:12.000769Z","shell.execute_reply.started":"2022-07-25T12:11:12.000583Z","shell.execute_reply":"2022-07-25T12:11:12.000607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Validate blended model\npredict_model(blender)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:11:12.001767Z","iopub.status.idle":"2022-07-25T12:11:12.002291Z","shell.execute_reply.started":"2022-07-25T12:11:12.002092Z","shell.execute_reply":"2022-07-25T12:11:12.002115Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Finalize blended model\nfinal_blender = finalize_model(blender)\nprint(final_blender)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:11:12.003307Z","iopub.status.idle":"2022-07-25T12:11:12.003841Z","shell.execute_reply.started":"2022-07-25T12:11:12.003592Z","shell.execute_reply":"2022-07-25T12:11:12.003622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Inference with blended model\ntest_blend_predictions = predict_model(final_blender, data=test)\nsubmission_df['SalePrice'] = test_blend_predictions['Label']\nsubmission_df.to_csv('submission_blender.csv', index=False)\ntest_blend_predictions","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:11:12.004855Z","iopub.status.idle":"2022-07-25T12:11:12.005186Z","shell.execute_reply.started":"2022-07-25T12:11:12.005011Z","shell.execute_reply":"2022-07-25T12:11:12.005035Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =\"5\"></a><h1 style=\"background:#05445E; border:0; border-radius: 12px; color:#D3D3D3\"><center>5. Deep Learning</center></h1>","metadata":{}},{"cell_type":"code","source":"## Import dependencies \nimport sklearn\nfrom sklearn.model_selection import KFold\n\nimport tensorflow as tf\nfrom tensorflow import keras\nfrom tensorflow.keras import layers\nimport tensorflow_addons as tfa","metadata":{"execution":{"iopub.status.busy":"2022-07-25T12:11:12.006550Z","iopub.status.idle":"2022-07-25T12:11:12.006883Z","shell.execute_reply.started":"2022-07-25T12:11:12.006704Z","shell.execute_reply":"2022-07-25T12:11:12.006733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =\"5.1\"></a><h2 style=\"background:#75E6DA; border:0; border-radius: 12px; color:black\"><center>5.1 Creating Dataset</center></h2>","metadata":{}},{"cell_type":"code","source":"## Fill NaN in numerical columns with its median\ntrain[numerical_features] = train[numerical_features].fillna(train[numerical_features].median()) \n\n## Fill NaN in categorical columns with its mode\ntrain[categorical_features] = train[categorical_features].fillna(train[categorical_features].mode().iloc[0])  \n\ntrain.info()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:11:12.008410Z","iopub.status.idle":"2022-07-25T12:11:12.008831Z","shell.execute_reply.started":"2022-07-25T12:11:12.008648Z","shell.execute_reply":"2022-07-25T12:11:12.008673Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Box-Cox Transformation on target ('SalePrice')\nauto_boxcox, best_lambda = sp.stats.boxcox(train[target], lmbda=None)\ntrain['target_transformed'] = auto_boxcox\n\n## Reverse Transformation\n#x = sp.special.inv_boxcox(auto_boxcox, best_lambda)\n\n## Box-Cox + Standardization\nbox_cox_mean = train['target_transformed'].mean()\nbox_cox_std = train['target_transformed'].std()\ntrain['target_transformed_standardized'] = (train['target_transformed'] - box_cox_mean) / box_cox_std","metadata":{"_kg_hide-output":false,"execution":{"iopub.status.busy":"2022-07-25T12:11:12.010234Z","iopub.status.idle":"2022-07-25T12:11:12.010606Z","shell.execute_reply.started":"2022-07-25T12:11:12.010390Z","shell.execute_reply":"2022-07-25T12:11:12.010413Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Train validation split\nn_splits = exp_config['n_splits']\nkf = KFold(n_splits=n_splits)\ntrain['k_folds'] = -1\nfor fold, (train_idx, valid_idx) in enumerate(kf.split(train)):\n    train['k_folds'][valid_idx] = fold\n    \nfor i in range(n_splits):\n    print(f\"fold {i}: {len(train.query('k_folds==@i'))} samples\")","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:11:12.011801Z","iopub.status.idle":"2022-07-25T12:11:12.012142Z","shell.execute_reply.started":"2022-07-25T12:11:12.011957Z","shell.execute_reply":"2022-07-25T12:11:12.011983Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Hold-out validation\nvalid_fold = train.query(f'k_folds == 0').reset_index(drop=True)\ntrain_fold = train.query(f'k_folds != 0').reset_index(drop=True)\nprint(len(train_fold), len(valid_fold))","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:11:12.013366Z","iopub.status.idle":"2022-07-25T12:11:12.013924Z","shell.execute_reply.started":"2022-07-25T12:11:12.013727Z","shell.execute_reply":"2022-07-25T12:11:12.013754Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def df_to_dataset(dataframe, target=None,\n                  shuffle=False, repeat=False,\n                  batch_size=5, drop_remainder=False):\n    df = dataframe.copy()\n    if target is not None:\n        labels = df.pop(target)\n        data = {key: value[:, tf.newaxis] for key, value in df.items()}\n        data = dict(data)\n        ds = tf.data.Dataset.from_tensor_slices((data, labels))\n    else:\n        data = {key: value[:, tf.newaxis] for key, value in df.items()}\n        data = dict(data)\n        ds = tf.data.Dataset.from_tensor_slices(data)\n    \n    if shuffle:\n        ds = ds.shuffle(buffer_size=len(df))\n    if repeat:\n        ds = ds.repeat()\n    ds = ds.batch(batch_size, drop_remainder=drop_remainder)\n    ds = ds.prefetch(batch_size)\n    return ds","metadata":{"execution":{"iopub.status.busy":"2022-07-25T12:11:12.014937Z","iopub.status.idle":"2022-07-25T12:11:12.015452Z","shell.execute_reply.started":"2022-07-25T12:11:12.015269Z","shell.execute_reply":"2022-07-25T12:11:12.015291Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Create datasets\nbatch_size = exp_config['batch_size']\ntarget = 'target_transformed_standardized'\n\ntrain_ds = df_to_dataset(train_fold, \n                         target=target,\n                         shuffle=True,\n                         repeat=False,\n                         batch_size=batch_size,\n                         drop_remainder=True)\n\nvalid_ds = df_to_dataset(valid_fold,\n                         target=target,\n                         shuffle=False,\n                         repeat=False,\n                         batch_size=batch_size,\n                         drop_remainder=False)\n\nexample = next(iter(train_ds))[0]\ninput_dtypes = {}\nfor key in example:\n    input_dtypes[key] = example[key].dtype\n    print(f'{key}, shape:{example[key].shape}, {example[key].dtype}')","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:11:12.016523Z","iopub.status.idle":"2022-07-25T12:11:12.017037Z","shell.execute_reply.started":"2022-07-25T12:11:12.016853Z","shell.execute_reply":"2022-07-25T12:11:12.016876Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =\"5.2\"></a><h2 style=\"background:#75E6DA; border:0; border-radius: 12px; color:black\"><center>5.2 Preprocessing Model</center></h2>","metadata":{}},{"cell_type":"code","source":"def create_preprocessing_model(numerical, categorical, input_dtypes, df):\n    ## Create input layers\n    preprocess_inputs = {}\n    for key in numerical:\n        preprocess_inputs[key] = layers.Input(shape=(1,),\n                                              dtype=input_dtypes[key])\n    for key in categorical:\n        preprocess_inputs[key] = layers.Input(shape=(1,),\n                                              dtype=input_dtypes[key])\n    \n    ## Create preprocess layers\n    normalize_layers = {}\n    for key in numerical:\n        normalize_layer = layers.Normalization(mean=df[key].mean(),\n                                               variance=df[key].var())\n        normalize_layers[key] = normalize_layer\n        \n    lookup_layers = {}\n    for key in categorical:\n        if input_dtypes[key] == tf.string:\n            lookup_layer = layers.StringLookup(vocabulary=df[key].unique(),\n                                               output_mode='int')\n        elif input_dtypes[key] == tf.int64:\n            lookup_layer = layers.IntegerLookup(vocabulary=df[key].unique(),\n                                                output_mode='int')\n        lookup_layers[key] = lookup_layer\n        \n    ## Create outputs\n    preprocess_outputs = {}\n    for key in preprocess_inputs:\n        if key in normalize_layers:\n            output = normalize_layers[key](preprocess_inputs[key])\n            preprocess_outputs[key] = output\n        elif key in lookup_layers:\n            output = lookup_layers[key](preprocess_inputs[key])\n            preprocess_outputs[key] = output\n            \n    ## Create model\n    preprocessing_model = tf.keras.Model(preprocess_inputs,\n                                         preprocess_outputs)\n    \n    return preprocessing_model, lookup_layers","metadata":{"_kg_hide-output":false,"execution":{"iopub.status.busy":"2022-07-25T12:11:12.018026Z","iopub.status.idle":"2022-07-25T12:11:12.018570Z","shell.execute_reply.started":"2022-07-25T12:11:12.018365Z","shell.execute_reply":"2022-07-25T12:11:12.018387Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Create preprocessing model\npreprocessing_model, lookup_layers = create_preprocessing_model(\n    numerical_features, categorical_features, \n    input_dtypes, train_fold)\n\n## Apply the preprocessing model in tf.data.Dataset.map\ntrain_ds = train_ds.map(lambda x, y: (preprocessing_model(x), y),\n                        num_parallel_calls=tf.data.AUTOTUNE)\nvalid_ds = valid_ds.map(lambda x, y: (preprocessing_model(x), y),\n                        num_parallel_calls=tf.data.AUTOTUNE)\n\n## Display a preprocessed input sample\nexample = next(train_ds.take(1).as_numpy_iterator())[0]\nfor key in example:\n    print(f'{key}, shape: {example[key].shape}, {example[key].dtype}')","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:11:12.019568Z","iopub.status.idle":"2022-07-25T12:11:12.020080Z","shell.execute_reply.started":"2022-07-25T12:11:12.019892Z","shell.execute_reply":"2022-07-25T12:11:12.019914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =\"5.3\"></a><h2 style=\"background:#75E6DA; border:0; border-radius: 12px; color:black\"><center>5.3 Training Model</center></h2>","metadata":{}},{"cell_type":"code","source":"def create_training_model(numerical, categorical, input_dtypes, df,\n                          lookup_layers, emb_dim=1,\n                          model_units=[128,], \n                          dropout_rates=[0.2,]):\n    ## Create input layers\n    model_inputs = {}\n    for key in numerical:\n        model_inputs[key] = layers.Input(shape=(1,), \n                                         dtype='float32')\n    for key in categorical:\n        model_inputs[key] = layers.Input(shape=(1,), \n                                         dtype='int64')\n    \n    features = []\n    \n    for key in model_inputs:\n        if key in numerical:\n            features.append(model_inputs[key])\n        elif key in categorical:\n            ## Create embedding layers\n            embedding = layers.Embedding(\n                input_dim=lookup_layers[key].vocabulary_size(),\n                output_dim=emb_dim)\n            encoded_categorical = embedding(model_inputs[key])\n            encoded_categorical = tf.squeeze(encoded_categorical, axis=1)\n            features.append(encoded_categorical)\n        \n    ## Concatenate all features\n    x = tf.concat(features, axis=1)\n    \n    for units, dropout_rate in zip(model_units, dropout_rates):\n        feedforward = keras.Sequential([\n            layers.Dense(units, use_bias=False),\n            layers.BatchNormalization(),\n            layers.ReLU(),\n            layers.Dropout(dropout_rate),\n        ])\n        x = feedforward(x)\n        \n    final_layer = layers.Dense(units=1, activation=None)\n    model_outputs = final_layer(x)\n    \n    training_model = tf.keras.Model(inputs=model_inputs,\n                                    outputs=model_outputs)\n    return training_model","metadata":{"execution":{"iopub.status.busy":"2022-07-25T12:11:12.021252Z","iopub.status.idle":"2022-07-25T12:11:12.021798Z","shell.execute_reply.started":"2022-07-25T12:11:12.021550Z","shell.execute_reply":"2022-07-25T12:11:12.021583Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Create training model\nemb_dim = model_config['emb_dim']\nmodel_units = model_config['model_units']\ndropout_rates = model_config['dropout_rates']\n\ntraining_model = create_training_model(numerical_features,\n                                       categorical_features,\n                                       input_dtypes,\n                                       train_fold,\n                                       lookup_layers,\n                                       emb_dim=emb_dim,\n                                       model_units=model_units, \n                                       dropout_rates=dropout_rates)\n\n## Model compile and build\nlr = exp_config['learning_rate']\noptimizer = keras.optimizers.Adam(learning_rate=lr)\nloss_fn = keras.losses.MeanSquaredError()\n\ntraining_model.compile(optimizer=optimizer,\n                       loss=loss_fn,\n                       metrics=[keras.metrics.mean_squared_error,])\n\ntraining_model.summary()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:11:12.022670Z","iopub.status.idle":"2022-07-25T12:11:12.023003Z","shell.execute_reply.started":"2022-07-25T12:11:12.022829Z","shell.execute_reply":"2022-07-25T12:11:12.022851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =\"5.4\"></a><h2 style=\"background:#75E6DA; border:0; border-radius: 12px; color:black\"><center>5.4 Model Training</center></h2>","metadata":{}},{"cell_type":"code","source":"## Settings for Training\nepochs = exp_config['train_epochs']\nbatch_size = exp_config['batch_size']\nsteps_per_epoch = len(train_fold)//batch_size \n\n## For saving the best model\ncheckpoint_filepath = exp_config['checkpoint_filepath']\nmodel_checkpoint_callback = tf.keras.callbacks.ModelCheckpoint(\n    filepath=checkpoint_filepath, \n    save_weights_only=True, \n    monitor='val_loss', \n    mode='min', \n    save_best_only=True)\n\n## For the adjustment of learning rate\nreduce_lr = tf.keras.callbacks.ReduceLROnPlateau(\n    monitor='val_loss',\n    factor=0.5,\n    patience=3,\n    cooldown=10,\n    min_lr=1e-5,\n    verbose=1)\n\n## Model training\nhistory = training_model.fit(train_ds,\n                             epochs=epochs,\n                             shuffle=True,\n                             validation_data=valid_ds,\n                             callbacks=[model_checkpoint_callback,\n                                        reduce_lr])\n\n## Load the best parameters\ntraining_model.load_weights(checkpoint_filepath)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:11:12.024403Z","iopub.status.idle":"2022-07-25T12:11:12.025074Z","shell.execute_reply.started":"2022-07-25T12:11:12.024865Z","shell.execute_reply":"2022-07-25T12:11:12.024890Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Plot the train and valid losses\ndef plot_history(hist, title=None, valid=True):\n    plt.figure(figsize=(7, 5))\n    plt.plot(np.array(hist.index), hist['loss'], label='Train Loss')\n    if valid:\n        plt.plot(np.array(hist.index), hist['val_loss'], label='Valid Loss')\n    plt.xlabel('Epoch')\n    plt.ylabel('Loss')\n    plt.legend()\n    plt.title(title)\n    plt.show()\n    \nhist = pd.DataFrame(history.history)\nplot_history(hist)","metadata":{"execution":{"iopub.status.busy":"2022-07-25T12:11:12.026044Z","iopub.status.idle":"2022-07-25T12:11:12.026369Z","shell.execute_reply.started":"2022-07-25T12:11:12.026194Z","shell.execute_reply":"2022-07-25T12:11:12.026218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =\"5.5\"></a><h2 style=\"background:#75E6DA; border:0; border-radius: 12px; color:black\"><center>5.5 Inference</center></h2>","metadata":{}},{"cell_type":"code","source":"## Inference_model = preprocessing_model + training_model\ninference_inputs = preprocessing_model.input\ninference_outputs = training_model(preprocessing_model(inference_inputs))\ninference_model = tf.keras.Model(inputs=inference_inputs,\n                                 outputs=inference_outputs)","metadata":{"execution":{"iopub.status.busy":"2022-07-25T12:11:12.027239Z","iopub.status.idle":"2022-07-25T12:11:12.027599Z","shell.execute_reply.started":"2022-07-25T12:11:12.027386Z","shell.execute_reply":"2022-07-25T12:11:12.027410Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Fill NaN in numerical columns with its median\ntest[numerical_features] = test[numerical_features].fillna(train[numerical_features].median()) \n## Fill NaN in categorical columns with its mode\ntest[categorical_features] = test[categorical_features].fillna(train[categorical_features].mode().iloc[0])  \n\n## Create test dataset\ntest_ds = df_to_dataset(test,\n                        target=None,\n                        shuffle=False,\n                        repeat=False,\n                        batch_size=batch_size,\n                        drop_remainder=False,)","metadata":{"execution":{"iopub.status.busy":"2022-07-25T12:11:12.029114Z","iopub.status.idle":"2022-07-25T12:11:12.029774Z","shell.execute_reply.started":"2022-07-25T12:11:12.029554Z","shell.execute_reply":"2022-07-25T12:11:12.029582Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"preds = inference_model.predict(test_ds)\npreds = np.squeeze(preds)\n\ndef reverse_transformation(preds,\n                           mean=box_cox_mean,\n                           std=box_cox_std,\n                           box_cox_lambda=best_lambda):\n    x = (preds * std) + mean\n    x = sp.special.inv_boxcox(x, box_cox_lambda)\n    return x \n\nprice_inference = reverse_transformation(preds)\nsubmission_df['SalePrice'] = price_inference\nsubmission_df.to_csv('submission_dnn.csv', index=False)\nsubmission_df.head()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-25T12:11:12.031013Z","iopub.status.idle":"2022-07-25T12:11:12.031336Z","shell.execute_reply.started":"2022-07-25T12:11:12.031164Z","shell.execute_reply":"2022-07-25T12:11:12.031186Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}