{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <center> A Neural Network Model for House Prices","metadata":{"_cell_guid":"c9d041f0-a7fe-46ea-8c84-821572ccd9a3","_uuid":"ca3198b1525214a53f33a933ddf0f22ac7963468"}},{"cell_type":"markdown","source":"The purpose of this notebook is to **build a model** (Deep Neural Network aka **DNN**) with **Tensorflow**. We will see the differents steps to do that. This notebook is split in several parts:\n\n- I.    Importation & Devices Available\n- II.   Outliers\n- III.  Preprocessing\n- IV.   DNNRegressor for **Contiunuous** features\n- V.    Predictions\n- VI.   Example with **Leaky Relu**\n- VII.  DNNRegressor for Continuous and **Categorial**\n- VIII. Predictions bis\n- IX.   Shallow Neural Network\n- X.    Conclusion\n\nWe will expose **3 models**:\n\n1. The first one will use just the **continuous features**, \n2. The second one we will add the **categorical features** and finally we will use a \n3. Neural Network with just one layer.\n\n* There are no tuning and we will **use DNNRegressor with *Relu* for all activations functions** and the number of units by layer are: [200, 100, 50, 25, 12]. So we have **5 layers**.\n\n* In the part VI I show how to **use another activation function** with the example of **Leaky Relu**. \n\n* Finally I will try to use a **Shallow Neural Network** (just with one Hidden Layer) just for fun.\n\nIf you have an idea to improve the performance of the model: Share it ! Fork it ! And play with it !","metadata":{}},{"cell_type":"markdown","source":"# INTRODUCTION\n\nYou have some experience with R or Python and machine learning basics. This is a perfect competition for data science students who have completed an online course in machine learning and are looking to expand their skill set before trying a featured competition\n\n**Competition Description**\n\nAsk a home buyer to describe their dream house, and they probably won't begin with the height of the basement ceiling or the proximity to an east-west railroad. But this playground competition's dataset proves that much more influences price negotiations than the number of bedrooms or a white-picket fence.\n\nWith 79 explanatory variables describing (almost) every aspect of residential homes in Ames, Iowa, this competition challenges you to predict the final price of each home.\n\n**Practice Skills**\n\n1. Creative **feature engineering** \n2. **Advanced regression techniques** like **random forest** and **gradient boosting**\n\n**Acknowledgments**\n\nThe [Ames Housing dataset](http://www.amstat.org/publications/jse/v19n3/decock.pdf) was compiled by Dean De Cock for use in data science education. It's an incredible alternative for data scientists looking for a modernized and expanded version of the often cited **Boston Housing dataset**. \n\nPhoto by [Tom Thain](https://unsplash.com/@tthfilms) on Unsplash.\n\n**File descriptions**\n\n**train.csv** - the training set \n\n**test.csv** - the test set\n\n**data_description.txt** - full description of each column, originally prepared by Dean De Cock but lightly edited to match the column names used here\n\n**sample_submission.csv** - a benchmark submission from a *linear regression* on year and month of sale, lot square footage, and number of bedrooms\n\n","metadata":{}},{"cell_type":"markdown","source":"# **Data fields**\n\nHere's a brief version of what you'll find in the data description file.\n\n**SalePrice** - the property's sale price in dollars. This is the target variable that you're trying to predict.\n\n**MSSubClass**: The building class\n\n**MSZoning**: The general zoning classification\n\n**LotFrontage**: Linear feet of street connected to property\n\n**LotArea**: Lot size in square feet\n\n**Street**: Type of road access\n\n**Alley**: Type of alley access\n\nLotShape: General shape of property\n\nLandContour: Flatness of the property\n\n**Utilities**: Type of utilities available\n\nLotConfig: Lot configuration\n\nLandSlope: Slope of property\n\nNeighborhood: Physical locations within Ames city limits\n\nCondition1: Proximity to main road or railroad\n\nCondition2: Proximity to main road or railroad (if a second is present)\n\nBldgType: Type of dwelling\n\nHouseStyle: Style of dwelling\n\nOverallQual: Overall material and finish quality\n\nOverallCond: Overall condition rating\n\nYearBuilt: Original construction date\n\nYearRemodAdd: Remodel date\n\nRoofStyle: Type of roof\n\nRoofMatl: Roof material\n\nExterior1st: Exterior covering on house\n\n**Exterior2nd**: Exterior covering on house (if more than one material)\n\n**MasVnrType**: Masonry veneer type\n\n**MasVnrArea**: Masonry veneer area in square feet\n\n**ExterQual**: Exterior material quality\n\n**ExterCond**: Present condition of the material on the exterior\n\n**Foundation**: Type of foundation\n\n**BsmtQual**: Height of the basement\n\n**BsmtCond**: General condition of the basement\n\n**BsmtExposure**: Walkout or garden level basement walls\n\n**BsmtFinType1**: Quality of basement finished area\n","metadata":{}},{"cell_type":"markdown","source":"**BsmtFinSF1**: Type 1 finished square feet\n\n**BsmtFinType2**: Quality of second finished area (if present)\n\n**BsmtFinSF2**: Type 2 finished square feet\n\n**BsmtUnfSF**: Unfinished square feet of basement area\n\n**TotalBsmtSF**: Total square feet of basement area\n\n**Heating**: Type of heating\n\n**HeatingQC**: Heating quality and condition\n\n**CentralAir**: Central air conditioning","metadata":{}},{"cell_type":"markdown","source":"Electrical: Electrical system\n\n1stFlrSF: First Floor square feet\n\n2ndFlrSF: Second floor square feet\n\nLowQualFinSF: Low quality finished square feet (all floors)\n\nGrLivArea: Above grade (ground) living area square feet\n\nBsmtFullBath: Basement full bathrooms\n\nBsmtHalfBath: Basement half bathrooms\n\nFullBath: Full bathrooms above grade\n\nHalfBath: Half baths above grade\n\nBedroom: Number of bedrooms above basement level\n\nKitchen: Number of kitchens\n\nKitchenQual: Kitchen quality\n\nTotRmsAbvGrd: Total rooms above grade (does not include bathrooms)\n\nFunctional: Home functionality rating\n\nFireplaces: Number of fireplaces\n\nFireplaceQu: Fireplace quality\n\nGarageType: Garage location\n\nGarageYrBlt: Year garage was built\n\nGarageFinish: Interior finish of the garage\n\nGarageCars: Size of garage in car capacity\n\nGarageArea: Size of garage in square feet\n\nGarageQual: Garage quality\n\nGarageCond: Garage condition\n\nPavedDrive: Paved driveway\n\nWoodDeckSF: Wood deck area in square feet\n\nOpenPorchSF: Open porch area in square feet\n\nEnclosedPorch: Enclosed porch area in square feet\n\n3SsnPorch: Three season porch area in square feet\n\nScreenPorch: Screen porch area in square feet\n\nPoolArea: Pool area in square feet\n\nPoolQC: Pool quality\n\nFence: Fence quality\n\nMiscFeature: Miscellaneous feature not covered in other categories\n\nMiscVal: $Value of miscellaneous feature\n\nMoSold: Month Sold\n\nYrSold: Year Sold\n\nSaleType: Type of sale\n\nSaleCondition: Condition of sale","metadata":{}},{"cell_type":"markdown","source":"# <center> I. Importation & Devices Available","metadata":{"_cell_guid":"78d12eeb-72ea-43b0-b503-c8b9842e7d7b","_uuid":"bd03c1594867b43b3f3494591b788d8eb9eab920"}},{"cell_type":"markdown","source":"Before the importation I prefer to check the devices available. Sometimes we can have a problems with your **GPU** for example. And if you want to have a good performance you must do use GPU and not CPU. In our example we have just a CPU but now you have the code to check if your devices is detected.","metadata":{"_cell_guid":"cb0e6d26-0b33-4f57-9aa7-91bdb6e9d953","_uuid":"ff96a752268379cf111e275fee40722ad6780f18"}},{"cell_type":"code","source":"import os\nimport tensorflow as tf\n\n# To disable all logging output from TensorFlow, set the following environment variable before launching Python\nos.environ['TF_CPP_MIN_LOG_LEVEL'] = \"99\"","metadata":{"_uuid":"14491f7cb884be2c1ac35e9a2aca023604d1aed0","execution":{"iopub.status.busy":"2022-08-01T09:23:56.248120Z","iopub.execute_input":"2022-08-01T09:23:56.248482Z","iopub.status.idle":"2022-08-01T09:23:58.789032Z","shell.execute_reply.started":"2022-08-01T09:23:56.248417Z","shell.execute_reply":"2022-08-01T09:23:58.788195Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tensorflow.python.client import device_lib\n# enables you to list the devices available in the local process.\n# https://stackoverflow.com/questions/38559755/how-to-get-current-available-gpus-in-tensorflow\n\ndevice_lib.list_local_devices()","metadata":{"_cell_guid":"90b67cbc-419a-4fc8-9a91-8d8dc7f34396","_uuid":"ff972448f235077543f3a0f6b9391bcea18329b6","execution":{"iopub.status.busy":"2022-08-01T09:24:03.371462Z","iopub.execute_input":"2022-08-01T09:24:03.371857Z","iopub.status.idle":"2022-08-01T09:24:08.208470Z","shell.execute_reply.started":"2022-08-01T09:24:03.371786Z","shell.execute_reply":"2022-08-01T09:24:08.207465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tensorflow.python.client import device_lib\n\ndef get_available_gpus():\n    local_device_protos = device_lib.list_local_devices()\n    return [x.name for x in local_device_protos if x.device_type == 'GPU']\n\nget_available_gpus()","metadata":{"execution":{"iopub.status.busy":"2022-08-01T09:24:24.847634Z","iopub.execute_input":"2022-08-01T09:24:24.847995Z","iopub.status.idle":"2022-08-01T09:24:24.859455Z","shell.execute_reply.started":"2022-08-01T09:24:24.847935Z","shell.execute_reply":"2022-08-01T09:24:24.858533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now that we have checked the devices available we will test them wth a simple computation. Here we have an example with the computation on the CPU. But you can split the computation on your gpu with '/gpu:0'. If you want more GPU you can do 'with tf.device('/gpu:1'): ', 'with tf.device('/gpu:2'): ' etc...\n\n**In our example I display the log information.**","metadata":{"_cell_guid":"8932987f-35e8-4769-9feb-ad9916bf0109","_uuid":"8f79a27f340c23cc74add1419bb1c98cab99b183"}},{"cell_type":"code","source":"# Test with a simple computation\nimport tensorflow as tf\n\ntf.Session()\n\nwith tf.device('/cpu:0'):\n    a = tf.constant([1.0, 2.0, 3.0, 4.0, 5.0, 6.0], shape=[2, 3])\n# If you have gpu you can try this line to compute b with your GPU\n#with tf.device('/gpu:0'):    \n    b = tf.constant([1.0, 2.0, 3.0, 4.0, 5.0, 6.0], shape=[3, 2])\n    \nc = tf.matmul(a, b)\n\n# Creates a session with log_device_placement set to True.\n\n# What does TF ConfigProto do?\n# At first, TensorFlow uses tf.ConfigProto() to configure the session. \n# It can also take in parameters when running tasks by setting environmental variable \n# CUDA_VISIBLE_DEVICES.\n\nsess = tf.Session(config=tf.ConfigProto(log_device_placement=True))\n\nprint(sess.run(c))\n\n# Runs the op.\n# Log information\noptions = tf.RunOptions(output_partition_graphs=True)\nmetadata = tf.RunMetadata()\n# trace each iteration, e.g. tensorboard > graphs > session runs; \n# metadata also stores information like run times, memory consumption, e.g.\n\nc_val = sess.run(c, options=options, run_metadata=metadata)\n\nprint(metadata.partition_graphs)\n\nsess.close()","metadata":{"_cell_guid":"e14879c3-a5bd-4aee-8e06-139f643b7d18","_uuid":"03e668187944a3aa412b3bba2e2a06edb5eb9e20","execution":{"iopub.status.busy":"2022-08-01T09:24:38.870580Z","iopub.execute_input":"2022-08-01T09:24:38.870939Z","iopub.status.idle":"2022-08-01T09:24:38.996208Z","shell.execute_reply.started":"2022-08-01T09:24:38.870875Z","shell.execute_reply":"2022-08-01T09:24:38.994253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In this tutorial our data is composed to **1460 rows** with **81 features**: \n\n* **38 continuous features** \n* **43 categorical features**. \n\nAs exposed in the introduction we will use onlly the **continuous** features to build our first model.\n\nHere the objective is to predict the House Prices. In this case we have a **regression model** to build.\nSo our first data we will contain 37 features to explain the 'SalePrice'. We can see the list of features that we will use to build our first model.","metadata":{"_cell_guid":"cdb671d1-2803-4290-8e8b-a54a00d21d44","_uuid":"348514ad519851c2dc7e51dc04acdbe7d391d6fc"}},{"cell_type":"markdown","source":"__future__ module is a built-in module in Python that is used to **inherit** new features that will be available in the new Python versions.. This module includes all the latest functions which were not present in the previous version in Python. And we can use this by importing the __future__ module\n\nhttps://www.geeksforgeeks.org/__future__-module-in-python/#:~:text=__future__%20module%20is,the%20__future__%20module.","metadata":{}},{"cell_type":"code","source":"import __future__\nprint(__future__.all_feature_names)\nprint(\"-----------------------------------------------------------------------------------------------\")\nprint(__future__.print_function)\nprint(\"-----------------------------------------------------------------------------------------------\")\nprint(__future__.division)","metadata":{"execution":{"iopub.status.busy":"2022-08-01T09:25:15.460566Z","iopub.execute_input":"2022-08-01T09:25:15.460953Z","iopub.status.idle":"2022-08-01T09:25:15.470278Z","shell.execute_reply.started":"2022-08-01T09:25:15.460875Z","shell.execute_reply":"2022-08-01T09:25:15.468893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"__future__.absolute_import","metadata":{"execution":{"iopub.status.busy":"2022-08-01T09:25:19.721892Z","iopub.execute_input":"2022-08-01T09:25:19.722243Z","iopub.status.idle":"2022-08-01T09:25:19.729342Z","shell.execute_reply.started":"2022-08-01T09:25:19.722177Z","shell.execute_reply":"2022-08-01T09:25:19.727896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"__future__","metadata":{"execution":{"iopub.status.busy":"2022-08-01T09:25:23.053013Z","iopub.execute_input":"2022-08-01T09:25:23.053412Z","iopub.status.idle":"2022-08-01T09:25:23.059558Z","shell.execute_reply.started":"2022-08-01T09:25:23.053331Z","shell.execute_reply":"2022-08-01T09:25:23.058546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**What is rcParams in PyLab?**\n\nrcParams is a matplotlib. RcParams object, it is a dictionary-like variable which store some rc settings in matplotlib\n\nhttps://www.tutorialexample.com/understand-matplotlib-rcparams-a-beginner-guide-matplotlib-tutorial/#:~:text=rcParams%20is%20a%20matplotlib.,some%20rc%20settings%20in%20matplotlib.\n\n**What does rcParams do in Python?**\n\nChanging the Defaults: rcParams\n\nEach time Matplotlib loads, it defines a **runtime configuration (rc)** containing the default styles for every plot element you create. This configuration can be adjusted at any time using the plt.\n\nhttps://jakevdp.github.io/PythonDataScienceHandbook/04.11-settings-and-stylesheets.html#:~:text=Changing%20the%20Defaults%3A%20rcParams,any%20time%20using%20the%20plt.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\ntrain = pd.read_csv('../input/train.csv')\ntrain.head()\nprint(train.columns)","metadata":{"execution":{"iopub.status.busy":"2022-08-01T09:25:33.125167Z","iopub.execute_input":"2022-08-01T09:25:33.125511Z","iopub.status.idle":"2022-08-01T09:25:33.187722Z","shell.execute_reply.started":"2022-08-01T09:25:33.125446Z","shell.execute_reply":"2022-08-01T09:25:33.186487Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from __future__ import absolute_import\n# https://stackoverflow.com/questions/33743880/what-does-from-future-import-absolute-import-actually-do\n\nfrom __future__ import division\n# https://peps.python.org/pep-0238/:~:text=The%20future%20division%20statement%2C%20spelled,make%20true%20division%20the%20default.\n    \nfrom __future__ import print_function\n\nimport itertools\n\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nfrom pylab import rcParams\nimport matplotlib\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import MinMaxScaler\n\ntf.logging.set_verbosity(tf.logging.INFO) \n# https://www.tensorflow.org/api_docs/python/tf/compat/v1/logging/set_verbosity\n\n# Sets the threshold for what messages will be logged\n# The logging level documentation page basically tells you: \n# If you set it to the level as displayed (INFO), \n# then TensorFlow will tell you all messages that have the label INFO (or more critical).\n\n# Say you would only be interested in WARN or ERROR, then you could similarly \n# set tf.logging.set_verbosity(tf.logging.WARN)\n\nsess = tf.InteractiveSession()\n# What does TF InteractiveSession () do?\n# The only difference between Session and an InteractiveSession is \n# that InteractiveSession makes itself the default session so that you can call run() or eval() \n# without explicitly calling the session.\n\n\ntrain = pd.read_csv('../input/train.csv')\nprint('Shape of the train data with all features:', train.shape)\ntrain = train.select_dtypes(exclude=['object'])\nprint(\"\")\nprint('Shape of the train data with numerical features:', train.shape)\ntrain.drop('Id',axis = 1, inplace = True)\ntrain.fillna(0,inplace=True)\n\ntest = pd.read_csv('../input/test.csv')\ntest = test.select_dtypes(exclude=['object'])\nID = test.Id\ntest.fillna(0,inplace=True)\ntest.drop('Id',axis = 1, inplace = True)\n\nprint(\"\")\nprint(\"List of features contained our dataset:\",list(train.columns))","metadata":{"_cell_guid":"9e15ad43-8a3a-4840-9c0d-afa2ba2bd148","_uuid":"16485c4c885b6416be79eade3cfff6345861a8bd","execution":{"iopub.status.busy":"2022-08-01T09:26:35.078556Z","iopub.execute_input":"2022-08-01T09:26:35.078942Z","iopub.status.idle":"2022-08-01T09:26:35.371276Z","shell.execute_reply.started":"2022-08-01T09:26:35.078850Z","shell.execute_reply":"2022-08-01T09:26:35.370175Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <center> II. Outliers","metadata":{"_cell_guid":"c675ba7f-3d0c-4641-b133-e80711a19cf1","_uuid":"6b90c9ce6240afc8a36106661afe100170ae738a"}},{"cell_type":"markdown","source":"In this small part we will isolate the outliers with an **IsolationForest** (http://scikit-learn.org/stable/modules/generated/sklearn.ensemble.IsolationForest.html). I tried with and without this step and I had a better performance removing these rows.\n\nI haven't analysed the **test set** but I suppose that our train set looks like more at our data test without these outliers.\n","metadata":{"_cell_guid":"7b049d56-52b0-4b6c-8c35-490a57eaea87","_uuid":"ae94ba5f35444c2b46bd5bec0acd4a45e084bfc8"}},{"cell_type":"markdown","source":"**What type of algorithm is Isolation Forest?**\n\nThe Isolation Forest algorithm is a fast tree-based algorithm for anomaly detection. The algorithm uses the concept of path lengths in binary search trees to assign anomaly scores to each point in a dataset\n\nhttps://towardsdatascience.com/how-to-perform-anomaly-detection-with-the-isolation-forest-algorithm-e8c8372520bc\n\nhttps://www.analyticsvidhya.com/blog/2021/07/anomaly-detection-using-isolation-forest-a-complete-guide/\n\nhttps://medium.com/analytics-vidhya/understanding-isolation-forest-technique-for-anomaly-detection-261c702fa92f\n\nhttps://towardsdatascience.com/anomaly-detection-with-isolation-forest-visualization-23cd75c281e2\n\nhttps://blog.paperspace.com/anomaly-detection-isolation-forest/","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import IsolationForest\n# Return the anomaly score of each sample using the IsolationForest algorithm\n# https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.IsolationForest.html\n# Using Isolation Forest, we can not only detect anomalies faster but we also require less memory \n# compared to other algorithms. Isolation Forest isolates anomalies in the data points instead of \n# profiling normal data points.\n# https://blog.paperspace.com/anomaly-detection-isolation-forest/#:~:text=Using%20Isolation%20Forest%2C%20we%20can,of%20profiling%20normal%20data%20points.\n\n\n\nclf = IsolationForest(max_samples = 100, random_state = 42)\nclf.fit(train)\ny_noano = clf.predict(train)\nprint(y_noano[:10])\nprint(len(y_noano))\nprint(\"-------------------------------------------------------\")\ny_noano = pd.DataFrame(y_noano, columns = ['Top'])\nprint(y_noano[:10])\nprint(\"-------------------------------------------------------\")\ny_noano[y_noano['Top'] == 1].index.values\nprint(y_noano[:10])\nprint(\"-------------------------------------------------------\")\n\ntrain = train.iloc[y_noano[y_noano['Top'] == 1].index.values]\nprint(train.head())\nprint(\"-------------------------------------------------------\")\ntrain.reset_index(drop = True, inplace = True)\nprint(train.head())\nprint(\"-------------------------------------------------------\")\nprint(\"-------------------------------------------------------\")\nprint(\"Number of Outliers:\", y_noano[y_noano['Top'] == -1].shape[0])\nprint(\"Number of rows without outliers:\", train.shape[0])","metadata":{"_cell_guid":"599c148c-bee5-424e-80cc-afc1d7b1e3fb","_uuid":"08498903317073ce7632c3fbff677ae842ca8389","execution":{"iopub.status.busy":"2022-08-01T09:26:52.241314Z","iopub.execute_input":"2022-08-01T09:26:52.241697Z","iopub.status.idle":"2022-08-01T09:26:52.723541Z","shell.execute_reply.started":"2022-08-01T09:26:52.241633Z","shell.execute_reply":"2022-08-01T09:26:52.722479Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"clf","metadata":{"execution":{"iopub.status.busy":"2022-08-01T09:27:10.630862Z","iopub.execute_input":"2022-08-01T09:27:10.631221Z","iopub.status.idle":"2022-08-01T09:27:10.638008Z","shell.execute_reply.started":"2022-08-01T09:27:10.631156Z","shell.execute_reply":"2022-08-01T09:27:10.636965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head(10)","metadata":{"_cell_guid":"b6056d39-23cd-4821-989a-e6769ad298d8","_uuid":"d80cb31a8c00f704879e1c4771f701dd5d8031d0","execution":{"iopub.status.busy":"2022-08-01T09:27:14.918812Z","iopub.execute_input":"2022-08-01T09:27:14.919199Z","iopub.status.idle":"2022-08-01T09:27:14.960943Z","shell.execute_reply.started":"2022-08-01T09:27:14.919116Z","shell.execute_reply":"2022-08-01T09:27:14.959852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_noano = clf.predict(train)\ny_noano[:500]","metadata":{"execution":{"iopub.status.busy":"2022-08-01T09:27:26.776622Z","iopub.execute_input":"2022-08-01T09:27:26.777008Z","iopub.status.idle":"2022-08-01T09:27:26.865872Z","shell.execute_reply.started":"2022-08-01T09:27:26.776942Z","shell.execute_reply":"2022-08-01T09:27:26.864862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.DataFrame(y_noano, columns = ['Top'])[:10]","metadata":{"execution":{"iopub.status.busy":"2022-08-01T09:27:36.399196Z","iopub.execute_input":"2022-08-01T09:27:36.399541Z","iopub.status.idle":"2022-08-01T09:27:36.410646Z","shell.execute_reply.started":"2022-08-01T09:27:36.399478Z","shell.execute_reply":"2022-08-01T09:27:36.409524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <center> III. Preprocessing","metadata":{"_cell_guid":"96ba460b-8ca5-4479-b9a8-94c04bd08b53","_uuid":"546316d0afa2fad5717bb15a11186d2229d38cd7"}},{"cell_type":"markdown","source":"To **rescale** our data we will use the fonction **MinMaxScaler** of Scikit-learn. I am wondering if it is not interesting to use the same MinMaxScaler for Train and Test !","metadata":{"_cell_guid":"445cfd64-2d2f-4218-afb6-b5870d849fb1","_uuid":"4b5d09fd48475d5707699f1725810c01811b61a8"}},{"cell_type":"code","source":"list(train.columns)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T12:08:08.017880Z","iopub.execute_input":"2022-07-31T12:08:08.018286Z","iopub.status.idle":"2022-07-31T12:08:08.029624Z","shell.execute_reply.started":"2022-07-31T12:08:08.018237Z","shell.execute_reply":"2022-07-31T12:08:08.028707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(list(train.columns))","metadata":{"execution":{"iopub.status.busy":"2022-08-01T09:27:49.215935Z","iopub.execute_input":"2022-08-01T09:27:49.216349Z","iopub.status.idle":"2022-08-01T09:27:49.222913Z","shell.execute_reply.started":"2022-08-01T09:27:49.216271Z","shell.execute_reply":"2022-08-01T09:27:49.221975Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mat_train = np.matrix(train)\nprint(mat_train)\nprint('-------------------------------------------------------')\nprint(\"the shape of the new Matrix is: \", mat_train.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-01T09:28:35.104993Z","iopub.execute_input":"2022-08-01T09:28:35.105395Z","iopub.status.idle":"2022-08-01T09:28:35.124520Z","shell.execute_reply.started":"2022-08-01T09:28:35.105340Z","shell.execute_reply":"2022-08-01T09:28:35.123034Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mat_y = np.array(train.SalePrice).reshape((1314,1))\nprint(mat_y[:20])\nprint('-----------------------------')\nprint(mat_y.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-01T09:29:03.657171Z","iopub.execute_input":"2022-08-01T09:29:03.657551Z","iopub.status.idle":"2022-08-01T09:29:03.667720Z","shell.execute_reply.started":"2022-08-01T09:29:03.657477Z","shell.execute_reply":"2022-08-01T09:29:03.666571Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"col_train = list(train.columns)\ncol_train","metadata":{"execution":{"iopub.status.busy":"2022-08-01T09:29:27.938451Z","iopub.execute_input":"2022-08-01T09:29:27.938833Z","iopub.status.idle":"2022-08-01T09:29:27.946175Z","shell.execute_reply.started":"2022-08-01T09:29:27.938736Z","shell.execute_reply":"2022-08-01T09:29:27.944896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mat_train = np.matrix(train)\nmat_train.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-01T09:29:41.483228Z","iopub.execute_input":"2022-08-01T09:29:41.483578Z","iopub.status.idle":"2022-08-01T09:29:41.491385Z","shell.execute_reply.started":"2022-08-01T09:29:41.483501Z","shell.execute_reply":"2022-08-01T09:29:41.490229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mat_train[:1]","metadata":{"execution":{"iopub.status.busy":"2022-08-01T09:29:46.400779Z","iopub.execute_input":"2022-08-01T09:29:46.401164Z","iopub.status.idle":"2022-08-01T09:29:46.409558Z","shell.execute_reply.started":"2022-08-01T09:29:46.401089Z","shell.execute_reply":"2022-08-01T09:29:46.408551Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import warnings\nwarnings.filterwarnings('ignore')\n\ncol_train = list(train.columns)\ncol_train_bis = list(train.columns)\n\ncol_train_bis.remove('SalePrice')\n\nmat_train = np.matrix(train)\nmat_test  = np.matrix(test)\nmat_new = np.matrix(train.drop('SalePrice',axis = 1))\nmat_y = np.array(train.SalePrice).reshape((1314,1))\n\nprepro_y = MinMaxScaler()\nprepro_y.fit(mat_y)\n\nprepro = MinMaxScaler() #preprocessing\nprepro.fit(mat_train)\n\nprepro_test = MinMaxScaler()\nprepro_test.fit(mat_new)\n\ntrain = pd.DataFrame(prepro.transform(mat_train),columns = col_train)\ntest  = pd.DataFrame(prepro_test.transform(mat_test),columns = col_train_bis)\n\ntrain.head()","metadata":{"_cell_guid":"941c0846-55bf-45bb-a49d-be8084b9c056","_uuid":"bdbe7b6ceb041aace9b18fe3849e568f3ad37f99","execution":{"iopub.status.busy":"2022-08-01T09:35:28.759151Z","iopub.execute_input":"2022-08-01T09:35:28.759531Z","iopub.status.idle":"2022-08-01T09:35:28.840783Z","shell.execute_reply.started":"2022-08-01T09:35:28.759462Z","shell.execute_reply":"2022-08-01T09:35:28.839754Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"col_train_bis #is equal to col_train without 'SalePrice' feature","metadata":{"execution":{"iopub.status.busy":"2022-08-01T09:35:34.532358Z","iopub.execute_input":"2022-08-01T09:35:34.532744Z","iopub.status.idle":"2022-08-01T09:35:34.539803Z","shell.execute_reply.started":"2022-08-01T09:35:34.532664Z","shell.execute_reply":"2022-08-01T09:35:34.538308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* **To use Tensorflow we need to transform our data (features) in a special format**. \n\n* As a reminder we have just the **continuous** features. So the first function used is: **tf.contrib.layers.real_valued_column**.\n\nMore information here: https://www.tensorflow.org/api_docs/python/tf/contrib/layers/real_valued_column\n\nThe others cells allowed to us to create a train set and test set with our training data set.\n\nhttps://docs.w3cub.com/tensorflow~python/tf/contrib/layers/real_valued_column","metadata":{"_cell_guid":"8343b430-aac9-4fb1-b752-cfa9f619a9c5","_uuid":"a24e009d75b66a0979ec4e87d34611b405ad8b85"}},{"cell_type":"code","source":"# FEATURES = col_train_bis\n# feature_cols = [tf.contrib.layers.real_valued_column(k) for k in FEATURES]\n# feature_cols","metadata":{"execution":{"iopub.status.busy":"2022-07-31T12:26:54.249094Z","iopub.execute_input":"2022-07-31T12:26:54.249382Z","iopub.status.idle":"2022-07-31T12:26:54.253346Z","shell.execute_reply.started":"2022-07-31T12:26:54.249328Z","shell.execute_reply":"2022-07-31T12:26:54.252373Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# COLUMNS = col_train\n# training_set = train[COLUMNS]\n# training_set.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T12:26:58.833450Z","iopub.execute_input":"2022-07-31T12:26:58.833799Z","iopub.status.idle":"2022-07-31T12:26:58.837221Z","shell.execute_reply.started":"2022-07-31T12:26:58.833738Z","shell.execute_reply":"2022-07-31T12:26:58.836394Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# prediction_set = train.SalePrice\n# prediction_set.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T12:27:12.659778Z","iopub.execute_input":"2022-07-31T12:27:12.660050Z","iopub.status.idle":"2022-07-31T12:27:12.665070Z","shell.execute_reply.started":"2022-07-31T12:27:12.660000Z","shell.execute_reply":"2022-07-31T12:27:12.663517Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# List of features\nCOLUMNS = col_train\nFEATURES = col_train_bis\nLABEL = \"SalePrice\"\n\n# Columns for tensorflow\nfeature_cols = [tf.contrib.layers.real_valued_column(k) for k in FEATURES]\n\n# Training set and Prediction set with the features to predict\ntraining_set = train[COLUMNS]\nprediction_set = train.SalePrice\n\n# Train and Test \nx_train, x_test, y_train, y_test = train_test_split(training_set[FEATURES] , prediction_set, test_size=0.33, random_state=42)\ny_train = pd.DataFrame(y_train, columns = [LABEL])\ntraining_set = pd.DataFrame(x_train, columns = FEATURES).merge(y_train, left_index = True, right_index = True)\ntraining_set.head()\n\n# Training for submission\ntraining_sub = training_set[col_train]\ntraining_set.head()","metadata":{"_cell_guid":"0b40bdb6-f813-4b94-94c6-599880f59445","_uuid":"091aee81e454d14b95d61cae7e35a90e4a376ec2","execution":{"iopub.status.busy":"2022-08-01T09:35:40.367624Z","iopub.execute_input":"2022-08-01T09:35:40.368053Z","iopub.status.idle":"2022-08-01T09:35:48.986890Z","shell.execute_reply.started":"2022-08-01T09:35:40.367975Z","shell.execute_reply":"2022-08-01T09:35:48.985804Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"training_sub.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-01T09:35:49.026240Z","iopub.execute_input":"2022-08-01T09:35:49.026765Z","iopub.status.idle":"2022-08-01T09:35:49.070831Z","shell.execute_reply.started":"2022-08-01T09:35:49.026681Z","shell.execute_reply":"2022-08-01T09:35:49.070006Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Same thing but for the test set\ny_test = pd.DataFrame(y_test, columns = [LABEL])\ntesting_set = pd.DataFrame(x_test, columns = FEATURES).merge(y_test, left_index = True, right_index = True)\ntesting_set.head()","metadata":{"_cell_guid":"0cd9a504-0fd9-47ca-a247-69a920911e35","_uuid":"9985b1ca4bfd4999989af752951676311c4f213e","execution":{"iopub.status.busy":"2022-08-01T09:36:18.441377Z","iopub.execute_input":"2022-08-01T09:36:18.441803Z","iopub.status.idle":"2022-08-01T09:36:18.487056Z","shell.execute_reply.started":"2022-08-01T09:36:18.441707Z","shell.execute_reply":"2022-08-01T09:36:18.486214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**What does TF contrib do?**\n\nIn general, tf.contrib contains contributed code. It is meant to contain features and contributions that eventually should get merged into core TensorFlow, but whose interfaces may still change, or which require some testing to see whether they can find broader acceptance.\n\nhttps://stackoverflow.com/questions/38060825/what-is-the-purpose-of-the-tf-contrib-module-in-tensorflow","metadata":{}},{"cell_type":"markdown","source":"# <center> IV. Deep Neural Network for continuous features","metadata":{"_cell_guid":"047f4ae6-7618-42d4-8ad0-4a19fd30ee81","_uuid":"44e0b00b9d8718876bbe814f68044918207faeab"}},{"cell_type":"markdown","source":"With **tf.contrib.learn** it is very easy to implement a Deep Neural Network. In our example we will have 5 hidden layers with repsectly 200, 100, 50, 25 and 12 units and the function of activation will be **Relu**.\n\n* The optimizer used in our case is an **Adagrad optimizer** (by default).\n\n* **TensorFlow's high-level machine learning API (tf. contrib.learn)** makes it easy to configure, train, and evaluate a variety of machine learning models\n\n**Very Important Tutorial:**\n\n**https://chromium.googlesource.com/external/github.com/tensorflow/tensorflow/+/r0.10/tensorflow/g3doc/tutorials/tflearn/index.md**\n\n* This package provides several ops that take care of creating variables that are used internally in a consistent way and provide the building blocks for many common machine learning algorithms.\n\nhttps://www.tensorflow.org/versions/r1.15/api_docs/python/tf/contrib/learn","metadata":{"_cell_guid":"0158e8c3-20e7-4a99-ab3a-b470ad4d0f9d","_uuid":"b166945d4a4af37865003549116816cf5bd71062"}},{"cell_type":"code","source":"# Model\ntf.logging.set_verbosity(tf.logging.ERROR)\nregressor = tf.contrib.learn.DNNRegressor(feature_columns=feature_cols, \n                                          activation_fn = tf.nn.relu, hidden_units=[200, 100, 50, 25, 12])#,\n                                         #optimizer = tf.train.GradientDescentOptimizer( learning_rate= 0.1 ))\n    \n#     https://www.tensorflow.org/versions/r1.15/api_docs/python/tf/contrib/learn/DNNRegressor","metadata":{"_cell_guid":"ed573d14-ffb1-4ae5-9f28-97057c20c6ce","_uuid":"6d2f589a8d2be2dd515929f9b9bdcf84736ac96e","execution":{"iopub.status.busy":"2022-08-01T09:46:26.598584Z","iopub.execute_input":"2022-08-01T09:46:26.598925Z","iopub.status.idle":"2022-08-01T09:46:26.610576Z","shell.execute_reply.started":"2022-08-01T09:46:26.598866Z","shell.execute_reply":"2022-08-01T09:46:26.609791Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Reset the index of training\ntraining_set.reset_index(drop = True, inplace =True)","metadata":{"_cell_guid":"cd61c1bb-0140-4b99-af7d-1feb99b43d9c","_uuid":"17cb0aa6571321dfd833e54de61fc08a27a724da","execution":{"iopub.status.busy":"2022-08-01T09:46:31.518116Z","iopub.execute_input":"2022-08-01T09:46:31.518468Z","iopub.status.idle":"2022-08-01T09:46:31.523966Z","shell.execute_reply.started":"2022-08-01T09:46:31.518389Z","shell.execute_reply":"2022-08-01T09:46:31.522658Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"FEATURES","metadata":{"execution":{"iopub.status.busy":"2022-08-01T09:38:59.348198Z","iopub.execute_input":"2022-08-01T09:38:59.348527Z","iopub.status.idle":"2022-08-01T09:38:59.357383Z","shell.execute_reply.started":"2022-08-01T09:38:59.348468Z","shell.execute_reply":"2022-08-01T09:38:59.356350Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def input_fn(data_set, pred = False):\n    \n    if pred == False:\n        \n        feature_cols = {k: tf.constant(data_set[k].values) for k in FEATURES}\n        labels = tf.constant(data_set[LABEL].values)\n        \n        return feature_cols, labels\n\n    if pred == True:\n        feature_cols = {k: tf.constant(data_set[k].values) for k in FEATURES}\n        \n        return feature_cols","metadata":{"_cell_guid":"4de80d3b-6f4d-4257-8b8c-763215173e28","_uuid":"7d121349efc922e5a69d3ef7e9aef69954e3b42a","execution":{"iopub.status.busy":"2022-08-01T09:46:36.523109Z","iopub.execute_input":"2022-08-01T09:46:36.523450Z","iopub.status.idle":"2022-08-01T09:46:36.538637Z","shell.execute_reply.started":"2022-08-01T09:46:36.523387Z","shell.execute_reply":"2022-08-01T09:46:36.537524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Deep Neural Network Regressor DNNRegressor with the training set which contain the data split by train test split\nregressor.fit(input_fn=lambda: input_fn(training_set), steps=2000)\n              \n# https://www.tensorflow.org/versions/r1.15/api_docs/python/tf/contrib/learn/DNNRegressor","metadata":{"_cell_guid":"aed9be24-12ad-48ca-b2eb-bf7c11d6a54a","_uuid":"8324e3670b648792b4010d495eb3564554077d80","execution":{"iopub.status.busy":"2022-08-01T09:47:01.288936Z","iopub.execute_input":"2022-08-01T09:47:01.289336Z","iopub.status.idle":"2022-08-01T09:47:08.797689Z","shell.execute_reply.started":"2022-08-01T09:47:01.289258Z","shell.execute_reply":"2022-08-01T09:47:08.796806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**x = lambda arguments : expression**\n\nLamda is just one line anonymous function\nUseful when writing **function inside function**\nit can take multiple arguments but computes only one expression\n\nexample:\n\nadd = lambda a, b : a + b\n\nadd(3,6) ## 9\n\nhttps://www.w3schools.com/python/python_lambda.asp\n\n**What is the lambda function in Python?**\nA lambda function is a small anonymous function. A lambda function can take any number of arguments, but can only have one expression\n\n**Why is lambda useful in Python?**\nThe lambda keyword in Python provides a shortcut for declaring small anonymous functions. Lambda functions behave just like regular functions declared with the def keyword. They can be used whenever function objects are required","metadata":{}},{"cell_type":"code","source":"# Evaluation on the test set created by train_test_split\nev = regressor.evaluate(input_fn=lambda: input_fn(testing_set), steps=1)","metadata":{"_cell_guid":"eb9cad74-6898-4bb5-a563-953a060f5448","_uuid":"3e21c5eaba39aa7ad09df92c051ef22c5c95fd0d","execution":{"iopub.status.busy":"2022-08-01T09:47:15.840647Z","iopub.execute_input":"2022-08-01T09:47:15.841049Z","iopub.status.idle":"2022-08-01T09:47:16.419083Z","shell.execute_reply.started":"2022-08-01T09:47:15.840973Z","shell.execute_reply":"2022-08-01T09:47:16.418230Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(regressor.get_params)\nprint(\"--------------------------------------------------------------------------------------\")\nprint(regressor.params)\nprint(\"--------------------------------------------------------------------------------------\")\nprint(regressor.predict_scores)\nprint(\"--------------------------------------------------------------------------------------\")","metadata":{"execution":{"iopub.status.busy":"2022-08-01T09:50:56.311640Z","iopub.execute_input":"2022-08-01T09:50:56.312009Z","iopub.status.idle":"2022-08-01T09:50:56.323199Z","shell.execute_reply.started":"2022-08-01T09:50:56.311946Z","shell.execute_reply":"2022-08-01T09:50:56.322111Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the score on the testing set\n# 0.002X in average\nloss_score1 = ev[\"loss\"]\nprint(\"Final Loss on the testing set: {0:f}\".format(loss_score1))","metadata":{"_cell_guid":"0c163468-e38f-4340-ac6e-198f25739d12","_uuid":"fe7df945092e0de175d64056b23dfd9da28747f9","execution":{"iopub.status.busy":"2022-08-01T09:53:26.426799Z","iopub.execute_input":"2022-08-01T09:53:26.427116Z","iopub.status.idle":"2022-08-01T09:53:26.433841Z","shell.execute_reply.started":"2022-08-01T09:53:26.427053Z","shell.execute_reply":"2022-08-01T09:53:26.432466Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"testing_set.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-01T09:54:29.673054Z","iopub.execute_input":"2022-08-01T09:54:29.673431Z","iopub.status.idle":"2022-08-01T09:54:29.680514Z","shell.execute_reply.started":"2022-08-01T09:54:29.673370Z","shell.execute_reply":"2022-08-01T09:54:29.679305Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**islice() function**\nThis iterator selectively prints the values mentioned in its iterable container passed as an argument.\n\nSyntax:\n\n**islice(iterable, start, stop, step)**\n\nhttps://www.geeksforgeeks.org/python-itertools-islice/","metadata":{}},{"cell_type":"code","source":"from itertools import islice\nfor i in islice(range(20), 5): \n    print(i)\n\nli = [2, 4, 5, 7, 8, 10, 20] \n  \n# Slicing the list\nprint(list(itertools.islice(li, 1, 6, 2)))","metadata":{"execution":{"iopub.status.busy":"2022-08-01T10:02:05.036229Z","iopub.execute_input":"2022-08-01T10:02:05.036552Z","iopub.status.idle":"2022-08-01T10:02:05.047943Z","shell.execute_reply.started":"2022-08-01T10:02:05.036493Z","shell.execute_reply":"2022-08-01T10:02:05.046650Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Predictions\ny = regressor.predict(input_fn=lambda: input_fn(testing_set))\npredictions = list(itertools.islice(y, testing_set.shape[0]))\n# Itertools islice method returns an iterator which returns the individual values on iterating or traversing over","metadata":{"_cell_guid":"c5c517b3-6011-4b05-9b34-ded41d82f8e4","_uuid":"6ba45c1abe834c716c5445dde3b794af05e1aeff","execution":{"iopub.status.busy":"2022-08-01T09:53:50.588951Z","iopub.execute_input":"2022-08-01T09:53:50.589331Z","iopub.status.idle":"2022-08-01T09:53:51.045004Z","shell.execute_reply.started":"2022-08-01T09:53:50.589254Z","shell.execute_reply":"2022-08-01T09:53:51.043901Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y","metadata":{"execution":{"iopub.status.busy":"2022-07-31T12:41:58.104643Z","iopub.execute_input":"2022-07-31T12:41:58.104916Z","iopub.status.idle":"2022-07-31T12:41:58.109459Z","shell.execute_reply.started":"2022-07-31T12:41:58.104863Z","shell.execute_reply":"2022-07-31T12:41:58.108631Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions[:5]","metadata":{"execution":{"iopub.status.busy":"2022-07-31T12:42:00.790148Z","iopub.execute_input":"2022-07-31T12:42:00.790417Z","iopub.status.idle":"2022-07-31T12:42:00.795722Z","shell.execute_reply.started":"2022-07-31T12:42:00.790366Z","shell.execute_reply":"2022-07-31T12:42:00.794837Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"testing_set.shape[0]","metadata":{"execution":{"iopub.status.busy":"2022-08-01T10:02:48.544312Z","iopub.execute_input":"2022-08-01T10:02:48.544666Z","iopub.status.idle":"2022-08-01T10:02:48.551440Z","shell.execute_reply.started":"2022-08-01T10:02:48.544586Z","shell.execute_reply":"2022-08-01T10:02:48.550333Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# list(itertools.islice(y, testing_set.shape[0]))","metadata":{"execution":{"iopub.status.busy":"2022-07-31T12:42:05.550310Z","iopub.execute_input":"2022-07-31T12:42:05.550615Z","iopub.status.idle":"2022-07-31T12:42:05.562539Z","shell.execute_reply.started":"2022-07-31T12:42:05.550539Z","shell.execute_reply":"2022-07-31T12:42:05.561631Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <center> V. Predictions and submission","metadata":{"_cell_guid":"87746ee4-64f8-49ba-9a23-85d21261dc6b","_uuid":"d75c2f98f4524273133765d146f131624ae7569c"}},{"cell_type":"markdown","source":"Let's go to prepare our first submission ! \n\n* Data Preprocessed: checked ! \n* Outlier excluded: checked ! \n* Model built: : checked!\n\nNext step: Used our model to make the predictions with the data set **Test**. And add one **graphic** to see the difference between the reality and the predictions.","metadata":{"_cell_guid":"1a08cbad-7b61-4c37-8686-5cc533d496ac","_uuid":"0cebdd59654ca804beddc4f1a590877e59aa4e8a"}},{"cell_type":"code","source":"prepro_y","metadata":{"execution":{"iopub.status.busy":"2022-08-01T10:03:10.620451Z","iopub.execute_input":"2022-08-01T10:03:10.620812Z","iopub.status.idle":"2022-08-01T10:03:10.627487Z","shell.execute_reply.started":"2022-08-01T10:03:10.620743Z","shell.execute_reply":"2022-08-01T10:03:10.625812Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Now we can apply the **inverse_transform** method to this set of test points. Each test point is a two dimensional point lying somewhere in the embedding space. \n\n* The inverse_transform method will convert this into an approximation of the high dimensional representation that would have been embedded into such a location","metadata":{}},{"cell_type":"code","source":"prepro_y.inverse_transform(np.array(predictions).reshape(434,1))[:10]","metadata":{"execution":{"iopub.status.busy":"2022-08-01T10:03:21.318858Z","iopub.execute_input":"2022-08-01T10:03:21.319185Z","iopub.status.idle":"2022-08-01T10:03:21.329150Z","shell.execute_reply.started":"2022-08-01T10:03:21.319125Z","shell.execute_reply":"2022-08-01T10:03:21.327891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions = pd.DataFrame(prepro_y.inverse_transform(np.array(predictions).reshape(434,1)),columns = ['Prediction'])","metadata":{"_cell_guid":"0fb8ed1b-c0df-41b9-a764-55d675e8aa7c","_uuid":"bf2281f14b4bcf56aac14ad691b1729e50cff5d8","execution":{"iopub.status.busy":"2022-08-01T10:03:36.885921Z","iopub.execute_input":"2022-08-01T10:03:36.886246Z","iopub.status.idle":"2022-08-01T10:03:36.892492Z","shell.execute_reply.started":"2022-08-01T10:03:36.886167Z","shell.execute_reply":"2022-08-01T10:03:36.891340Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions[:10]","metadata":{"execution":{"iopub.status.busy":"2022-08-01T10:03:44.602400Z","iopub.execute_input":"2022-08-01T10:03:44.602721Z","iopub.status.idle":"2022-08-01T10:03:44.615589Z","shell.execute_reply.started":"2022-08-01T10:03:44.602652Z","shell.execute_reply":"2022-08-01T10:03:44.614493Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(len(COLUMNS))\nprint(COLUMNS)","metadata":{"execution":{"iopub.status.busy":"2022-08-01T10:04:48.269617Z","iopub.execute_input":"2022-08-01T10:04:48.269946Z","iopub.status.idle":"2022-08-01T10:04:48.276139Z","shell.execute_reply.started":"2022-08-01T10:04:48.269884Z","shell.execute_reply":"2022-08-01T10:04:48.275051Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"testing_set.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-01T10:03:48.471297Z","iopub.execute_input":"2022-08-01T10:03:48.471633Z","iopub.status.idle":"2022-08-01T10:03:48.481100Z","shell.execute_reply.started":"2022-08-01T10:03:48.471575Z","shell.execute_reply":"2022-08-01T10:03:48.479827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"reality = pd.DataFrame(prepro.inverse_transform(testing_set), columns = [COLUMNS]).SalePrice","metadata":{"_cell_guid":"e4e48b49-f405-4e57-b6cb-587607253aab","_uuid":"aa904a7f562129ac5f362e39089e6923f78d1acc","execution":{"iopub.status.busy":"2022-08-01T10:05:59.303379Z","iopub.execute_input":"2022-08-01T10:05:59.303815Z","iopub.status.idle":"2022-08-01T10:05:59.317271Z","shell.execute_reply.started":"2022-08-01T10:05:59.303756Z","shell.execute_reply":"2022-08-01T10:05:59.316262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"reality[:10]","metadata":{"execution":{"iopub.status.busy":"2022-08-01T10:06:12.016439Z","iopub.execute_input":"2022-08-01T10:06:12.016793Z","iopub.status.idle":"2022-08-01T10:06:12.030126Z","shell.execute_reply.started":"2022-08-01T10:06:12.016702Z","shell.execute_reply":"2022-08-01T10:06:12.028914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"matplotlib.rc('xtick', labelsize=30) \nmatplotlib.rc('ytick', labelsize=30) \n\nfig, ax = plt.subplots(figsize=(50, 40))\n\nplt.style.use('ggplot')\nplt.plot(predictions.values, reality.values, 'ro')\nplt.xlabel('Predictions', fontsize = 30)\nplt.ylabel('Reality', fontsize = 30)\nplt.title('Predictions x Reality on dataset Test', fontsize = 30)\nax.plot([reality.min(), reality.max()], [reality.min(), reality.max()], 'k--', lw=4)\nplt.show()","metadata":{"_cell_guid":"450559b2-bff3-427f-8dec-569191bd8c94","_uuid":"a11028c04b0a6981169196e2c6fe84c245799b8e","execution":{"iopub.status.busy":"2022-08-01T10:06:35.803996Z","iopub.execute_input":"2022-08-01T10:06:35.804457Z","iopub.status.idle":"2022-08-01T10:06:37.077252Z","shell.execute_reply.started":"2022-08-01T10:06:35.804358Z","shell.execute_reply":"2022-08-01T10:06:37.076245Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_predict = regressor.predict(input_fn=lambda: input_fn(test, pred = True))\n\ndef to_submit(pred_y,name_out):\n    y_predict = list(itertools.islice(pred_y, test.shape[0]))\n    y_predict = pd.DataFrame(prepro_y.inverse_transform(np.array(y_predict).reshape(len(y_predict),1)), columns = ['SalePrice'])\n    y_predict = y_predict.join(ID)\n    y_predict.to_csv(name_out + '.csv',index=False)\n    \nto_submit(y_predict, \"submission_continuous\")","metadata":{"_cell_guid":"6008e5e1-29d7-42ac-b81a-d9e761a7e956","_uuid":"3be76f4c812b3676a6816d8eb9353f049948ac49","execution":{"iopub.status.busy":"2022-08-01T10:10:14.525948Z","iopub.execute_input":"2022-08-01T10:10:14.526297Z","iopub.status.idle":"2022-08-01T10:10:15.005519Z","shell.execute_reply.started":"2022-08-01T10:10:14.526220Z","shell.execute_reply":"2022-08-01T10:10:15.004660Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_predict","metadata":{"execution":{"iopub.status.busy":"2022-08-01T10:10:30.620722Z","iopub.execute_input":"2022-08-01T10:10:30.621080Z","iopub.status.idle":"2022-08-01T10:10:30.627783Z","shell.execute_reply.started":"2022-08-01T10:10:30.621020Z","shell.execute_reply":"2022-08-01T10:10:30.626446Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <center> VI. Leaky Relu","metadata":{"_cell_guid":"75a14830-c565-4efa-a298-e45ebf11055a","_uuid":"b4ac8e0fa1f7ea486880e006a5a70b37f91dd44a"}},{"cell_type":"markdown","source":"An example with another activation function: **Leaky Relu** ! We can create this new function with Relu. As a reminder Relu is Max(x,0) and Leaky Relu is the function Max(x, delta*x). In our case we can take delta = 0.01\n\nhttps://machinelearningmastery.com/choose-an-activation-function-for-deep-learning/#:~:text=for%20Output%20Layers-,Activation%20Functions,a%20layer%20of%20the%20network.\n\n**Activation Functions**\nAn activation function in a neural network defines how the weighted sum of the input is transformed into an output from a node or nodes in a layer of the network\n\n**For more info about AF and diving in more references, here below you can find more links :**\n\n* https://en.wikipedia.org/wiki/Activation_function\n\n* https://towardsdatascience.com/activation-functions-neural-networks-1cbd9f8d91d6\n\n* https://www.geeksforgeeks.org/activation-functions-neural-networks/\n\n* https://ml-cheatsheet.readthedocs.io/en/latest/activation_functions.html\n\n* https://www.analyticsvidhya.com/blog/2020/01/fundamentals-deep-learning-activation-functions-when-to-use-them/\n\n* https://keras.io/api/layers/activations/\n\n* https://medium.com/analytics-vidhya/activation-functions-all-you-need-to-know-355a850d025e\n\n* https://medium.com/@snaily16/what-why-and-which-activation-functions-b2bf748c0441\n\n* https://towardsdatascience.com/activation-functions-neural-networks-1cbd9f8d91d6\n\n* https://towardsdatascience.com/activation-functions-you-might-have-missed-79d72fc080a5\n\n* https://medium.datadriveninvestor.com/deep-learning-best-practices-activation-functions-weight-initialization-methods-part-1-c235ff976ed\n\n* https://www.turing.com/kb/how-to-choose-an-activation-function-for-deep-learning","metadata":{}},{"cell_type":"markdown","source":"![image.png](attachment:23c65332-8f96-475e-9dac-af932859424b.png)","metadata":{},"attachments":{"23c65332-8f96-475e-9dac-af932859424b.png":{"image/png":"iVBORw0KGgoAAAANSUhEUgAAAhcAAAELCAIAAADlRzNiAAAgAElEQVR4nO3df3Ab9Z038E+WzbIII4QIrnFcxzhuCCY1wTZpaE0JxUCgpvBckzj3TEin3D11j/IMf4Gnk+ncMX0yTOjcTJk+UAxc26Nmnovr3OSoL02JoaF1ORpskQuOS4OjOqojHGMUVfgUnapsnj++ymYjy9JK2l19JL1ff3TW+vnl3Y0+393v7ve75Ny5cwQAAJAXqdgNAACAEoYqAgAA+UMVAQCA/KGKAABA/lBFAAAgf6giAACQP1QRAADIn+zkl8ViMeOfkiRpmuZkA0oXsjIPWZkkSRIRISszkJWiKCKEhRytIqFQyPin1+tNeYQDVVVTql3RqarqcrmQlRmqqiqKEolEit2Qi8iyLMsyt6zcbjcRMcyKiBKJRLEbchGeWamqmkgknMnK6/Wqqpr2KZzRAnCIhnkiTENWJQRVBMBe+g+itGRJcVvCH7Iyj0+hRRUBsJf+g8jnnz1byMo8PlmhigDYC/1r85CVeXyyQhUBsBefPiN/yMo8PllZf43WxMTE6OgoEdXW1nZ2dlr++QClhU+fkT9kZR6frCw+FgmFQkNDQ52dnV1dXX6/3+fzWfv5ACWHT5+RP2RlHp+sLK4iExMTra2ttbW1Xq+3ubk5GAxa+/kAJYdPn5E/ZGUen6wsPqPV0dEhNqLRqN/vX7NmDRGFQqFwOExEqqoa71uRJGmx21iKSJZlbq2SZRlZmYSszBO3InNrlSDuPeSDZ1YiJcuzeuGFf3Gp8rbtm8w2w9qvFyYmJoaHh5ubm1tbW4koHA77/X4iWn3DDYqi6C+TJMn4JxOSJC12o3+xiCYhKzMkSZJlGVmZIX59GGbFEM+s7Nippv449ceq+zTJ/d3nDv7NA6sarmvQv2uxt1hfRYaGhqLR6LZt27xer3iksbGxsbGRiILBoHEKAVmWuc0oQFxn9ZAkCVmZgRlQzOM5qwdmQDHPjhlQfjx4RPvURiKKqzXeq736f7L+e76QxVXE5/NFo9EtW7ZY+7EApavog58lBFmZZ0dWr+0/GPnURrF9vTxO1GDmXRZXkVAodPTo0b//+78Xf95222242BcqnPFamqIPhDKHrMyzI6t3Z5eRm4jIfWr/pp4uk++yuIp0dnaibAAY8bmWhj9kZZ7lWfW9tDe6bCsRSVrk9pvd5t/Ia7ivFHV3dy9fYGRkhIiWL1++c+fOtG9JeXznzp3Lly93qMXgLD7X9fOHrMyzNqvpwPRsVfIKW8/sG+3rWs2/l9fldCXqkUce2bFjR7FbAUyhf20esjLP2qz27BtP1DxIRHJs+qv3rcnpvTgWAbAX+tfmISvzLMxq9JAvVPOg2F4RH6mrr8vp7TgWSQqEpF8fk4noEmnJWS3LVeHb1scdaRSUA/SvzUNW5lmY1a/G4+QlInJFRrdvfzDXt6OKJP36mPzBjDgyO5frIdrQ0NCRI0eMj+zevdu6pkFpw3VH5iEr86zKanBgaN6bvE395uq5PD4BVcQC9fX1LS0tRCTLMre7paDo0L82D1mZZ1VWf0isIYWIyH1q/909G/L4BFSRpKZq7cRcnqNEra2tvb291rYHygb61+YhK/MsyerZH74av3YLEUlapPvepvw+BFUkaX1jYn1jgvKa1cOS44/6+vrCPwQYQv/aPGRlXuFZTQem587fqb7s1P66+q/k9zmoIhaIRCLiBhFdfX29qAopT61Zs8bj8bS0tPT3999///3iPFggEBgaGurqMnunKJQW9K/NQ1bmFZ7Vnn3j2vmre7/1d3mWEEIVsUR/f39/f7/xEf0OkpSndu/e3dHRIZ669957RRU5cuRIV1fXt771LWdbDQ5B/9o8ZGVegVkZr+69QXqbKP9eLKpIoTJcjnXy5MnFntqxY8dDDz0UCATIcOACZQn9a/OQlXkFZvX6+4qYMqsqdHDTwwWdCEEVKRoUjwqB/rV5yMq8QrJ6+eW9UfdWsX3HmkIXTcG96wD2wv3Y5iEr8wrJ6oSSnDLLO7M3pymz0rekwPcDQGboX5uHrMzLO6tnnt+fUOuISNIiuU6ZlRaqCIC90L82D1mZl19Wo4d84eovie2a0L5cp8xK35LCPwIAMkD/2jxkZV5+Wb35bkST3ETkioz2/G3OU2alhSoCYC/0r81DVublkdXgwJC+IO5Kdcqyllj1QQCQFvrX5iEr8/LI6nisQWy4T+3ftMWy25xRRQqVYUHDlAUQu7u7U6b+NfNR3d3dTz75pPERLIxYWtC/Ng9ZmZdrVn0v7Y262yn3BXGzt8TCz4IUjzzyyMnzfvGLX0QikbT1Bsob+tfmISvzcspqOjA9471PbOe6IG5WqCIOaWlp6erq0ufUCgQC/f39Q0ND4XC4uA0Du6F/bR6yMi+nrPbsGxeD6nJs+rFvbrS2Jbh3Pen/ju359w/eMvniX2z9xzy+IhwOi5vV+/v7e3t7Ozo6IpHI448/vnv3bjGhFpQl9K/NQ1bmmc8qZUFcImsuzdKhiiTNzIdI06z9zCNHjhhPYQ0NDfX19YXD4d7eXjEtIxH19vbu3LkTayOWMcwNZR6yMs98VgUuiJsVqoiNgsGgvh2JRMTG0NCQOCIRZ7fq6+tT5gOGMoP+tXnIyjyTWRkXxL1zddyOlqCKJP2f2/+X2MhjlarFbNy4UUwCL/T29j777LP19fXhcPgHP/iB/rg4KIFyhf61ecjKPJNZ/V5bLza8M3vbH7R4RERAFXGO2+2empq67bbbGhoa9FNY4XB4fHw814/CZMAlBP1r85CVeWayevaHrybOL4hryZRZaaGKWCDtgoaLvbirq2vnzp39/f3btm0joqeeeiocDuuHIws/qqWl5Sc/+clf//Vfi8qBhRFLDvrX5iEr87JmZdWCuFmhilgg7YKGC1+2YsWK5557LhAI7NixY+fOna+88oq4zNc4tL7wo3bs2CHL8q233oqFEUsU+tfmISvzsma1+xeT2qeaiEiJvl/IgrhZLTnn4HXZxtFmIvJ6vaFQyLFvN8nCcZEM9BNZZgZFVFWdnZ0V972zWtvKmaxyoqqqoij6tQxMyLIsy3L0zBlWP45ut5sM130wIcsyEcX/8hdklZWqqolEIm1Wr+0/+NtY8kDks/HBwuc78Xq9qqqmfcrRYxGXy3XRd8tyyiMcSJLkQKtcLldtba3JF8uy3NTUVFNTY2uT8uBMVjkRv9fcWkVEkiRVXX55sVtxEfF7zTArIlKUQhfgsxbPrCRJkiQpbVbvzi7TF8Td/ugWW5vhaBWJRqPGP1VVTXmEA579a1mWkZUZorvELStR27hlJX4ZGWZFRIlEotgNuQjPrMSxyMKsUhbEtaTZix2IEGZAAbAbJvMwD1mZt1hW04FpaxfEzQpVBMBemBvKPGRl3mJZ7dk3LhbElWPT9l3de1FLHPgOgEqG647MQ1bmpc3KuCBu9fyIJQviZoUqAmAv9K/NQ1bmpc3qV+NxyxfEzd4SZ76mjFm1GpX4qJTHsSBVGUD/2jxkZd7CrAYHhua9G8T2zdVzjrUEVcQCWI0KMkD/2jxkZd7CrP6QSI6CuE/tv3vjBuda4tg3VYiU1agIC1JVPPSvzUNW5qVk1ffS3rhrNRFJWqT73iYnW4IZUJJmD01OvPhLIpJJTlCWa9U3vJhpAhJ9NSrCglSAuaFygazMM2YV/NNJ44K4dfW2zN27GFSRpIkXfxk8OJHfe9OuRkVEWJAKCP3rXCAr84xZ7dk3rtU8SPYsiJsVqogFAoGAvh2JRMRtroQFqYCI0L/OBbIyT8/HuCDuDdLbRE5P+I0qklRVe3Xe7+3q6lq4GlVfX9+JEyewIBWgf20essqJmPvk9XFZXxB30/YirBmBKpLU/uTW9ie3khVzQ4nVqIhoxYoVWJAK0L82D1nlRJbl/pcHozYviJsVrtGyUVdX19TUlH4W66mnnvrpT38qtsVqVDpx+VZLS0t/f79+rwkWpCoP6F+bh6xyddGCuPZPmZUWjkWsJ1ajOnLkSEtLy2ILUqVd2EqcFrv33nuxIFU5Qf/aPGSVk13f35uoeZBsXhA3K6xSlcry2c5zWpCKiAKBgBiu1xekUlXV5XJVQlaFY7hKlXbunLJ0KcOZ4RmuvCSyIn4zwzPMajow/U+j1WK+k9q5f7F7vhMuq1RVJo/Hk9OgOqvVDKFw6F+bh6zMMy6I69iUWWlhXATAXjjXbx6yMmlwYCjyqeR9IdfLOV+zYy1UEQB7YW4o85CVScdjDWLDfWp/4WuqFwhVBMBe6F+bh6zM6Htpb9TdTkSSFrmjzVPs5qCKANgM/WvzkFVW04Hp2arkOKtn9o1169uL2x5CFQGwG/rX5iGrrIwL4nY/uLbYzSFCFbFKOBweGRkxTqgFIKB/bR6yysw4ZdaK+EhDQwMxyApVpFDhcLi7u/vGG2/s7u6+9dZbb731Vv3mcwuXQfzyl7+MZRBLFPrX5iGrzH41npzjxBUZ3fbQA2K76FmhihSqp6fH4/EcPXr05MmTR48e/eIXv9jT06M/i2UQAf1r85BVBikL4vLJClWkUCMjIw899JDH4yEij8fz7W9/OxAIGNc61GEZxMqE/rV5yCoDfUFc78zeuzdu4JMV7l1POrTvu/53B02+eOuO/9S33W73z3/+8zVr1uiF5OTJk4u9EcsgViDcj20eslrMsz98NX7tFjJMmcUnK1SRpGhkRtPyeeP3vve9xx9/vL+/v6WlpaOj4/777zcWAyyDCHz6jPwhq7SmA9Nz5+9UX3Zqf139V4hTVqgiherq6hLnqd58882RkZHnnnuuo6Ojr69PHJpgGUTg02fkD1mlZVwQ91t/9xXxIJ+sUEWIiLRz5z7b8Td5vFHM19txHhEdOXKku7v7lVdeETO6YxlE4NNn5A9ZLfTa/oNpF8Tlk5VdVWR4eLizs9OmD7ectGTJ1XWtG7a2Uo6znY+Pj3d3dxsHQlpaWlpaWhYbLccyiBWIT5+RP2S10Luzy8hNRFQVOrjp4QtTZvHJyvprtPx+/8DAQMpSIuVqzZo1bre7t7dXf0SsXXj77bdnfmOuyyC2t7djGcQSxafPyB+ySvHyy8kps4jojjWK8Sk+WVl/LOL3+10uF7cFeWzi8Xh2797d09OzfPnyjo4OscDUI488stjpqfyWQaypqXnyySfn5+exDGIp4tNn5A9ZpTihJH9JvDN72x/caHyKT1a2rHXo9/tHRka2b9+u/+n3+4lo9Q03eK68Un8Zw5XyiEiSJC33q7UOHz4sisHam282/jdmFv7znw+/+y4RbdiwIXOTFEWJxWJTf5yaOjFFRA0rGhqua8i1kZbLLytbSZIky3I8Hi92Q1IxzEpRFCJimBVDRcnqu08PhM4viPu/b59f+E/eyZ1KURQRwkIYXbfG2rX5TIvmufLKzPVDJ0kSETVcx6J4QB4YVhG2kBURvTXydrj6S2K7JrSv4brtaV/GISsnqkhjY2NjYyMRBYNB48LFsiyzWsdYYHiEpKqqJEnIygyG664TkSzLWHfdJHE1PNZd/+V/zGqfaiUiV2S0528fTPvVqqomEglnsvJ6vYs9hRlQABxS9PmOSkiFZ2VcEHelOpX5xUXPClUEwF58rqXhD1kJZhbE5ZOVLWe09FNYAMDnWhr+kBWJBXGXbSUiSYt039u02Mv4ZIVjEQB78ekz8oespgPTM977xLZn9o26+rrFXsknK1QRAHvxWQeCP2S1Z9+4JrmJSI5NP/bNjRleyScrVBEAe/HpM/JX4VmlLIib+cV8skIVAbAXnz4jfxWe1evvJ2/rc0VGt29/MPOL+WSFKgJgLz59Rv4qOavBgSF9yqw7V2e/SZ5PVqgiAPbi02fkr5Kz+r22Xmx4Z/a2r2vN+no+WaGKANiLT5+Rv4rN6tkfvppQ68iwIG5WfLJCFQGwF58+I3+VmdWCBXEXvbrXiE9WqCIA9uLTZ+SvMrPa/YtJcXWvEn1fXxA3Kz5ZoYoA2ItPn5G/Cszqtf0H9SmzrpdzWPCUT1aoIgD24tNn5K8Cs3p3dpnYqAodXGzKrLT4ZIUqAmAvPn1G/iotq76XkgviSlokZUHcrPhkhSoCYC8+fUb+Kiqr6cD0bFVyQVzP7Btmru414pMVqgiAvfj0GfmrqKz27BsXV/fKsWmTV/ca8ckKVQTAXnz6jPxVTlajh3z6grjV8yMmr+414pMVqgiAvfj0GfmrnKx+NR4XV/eKBXHz+AQ+WaGKANiLT5+RvwrJanBgaN67QWzfXD2X34fwyQpVBMBefPqM/FVIVn9IJEdB3Kf2371xQ34fwicrVBEAe/HpM/JXCVn1vbQ37lpN2RbEzYpPVqgiAPbi02fkr+yzMr8gblZ8spIzPx0KhXw+34EDB8Sfd911V2dnp/2tAigffPqM/JV9Vnv2jWs1D5KJBXGz4pPVosciw8PDmzdvvvrqq3ft2uXz+YjI5/P19vYuWbKkp6fH7/c72EiAEsanz8hfeWdlXBD3BuntAj+NT1bpj0U2b94s/vdnP/vZwmdfeOGFnp6e1tbWXbt22ds6gNLHp8/IX3ln9fr7CrmJiFyR0U3bc5gyKy0+WaWvIj09PRnOXH3jG9/4xje+MTw8bFurAMqHsc9Y9H/wzJVxVi+/vDfq3iq2zSyImxWfrNKf0RIlxO/3h0KhlKfE2S39NQCQGZ8+I39lnNUJJTlllskFcbPik1Wma7T8fv8tt9xiPObo7e1ta2uzv1UA5YPP+Wv+yjWrZ57fn+uCuFnxySpTFens7Hzqqae++tWv9vb2+ny+tra24eHhsbExxxoHUAb49Bn5K8uspgPT+pRZNaF9hVzda8Qnqyz3i2zZsuWPf/zj888/39bW1tnZOTY21tpqwbEYQOXg02fkryyzMi6Im9+UWWnxySpLFfH5fHfddde6dev6+vqef/753t7ehSMlAJABnz4jf+WX1eDAUH4L4mbFJ6tMdx2Ks1hPPPGEuKJ3y5YtmzdvvuWWW44fP57fl6mqavxTkqSURzhg2CpZlhm2ipBVLhi2SpIkWvCvkglZznJDtMPyzup4rIEUIiL3qf3bHttkeatkWS56Vpm+3uPxGE9heTyeAwcOPP3003l/WSwWM/7pcrlSHikuccGcqqp8WqU3SdM0Pq0i3llJksStVeKfOp9WCYqi0IJ/lUWkZ0VEiUSi2M25SE5Z6Zfe9r20N7psKxFJWuT2m93WRq2dO+e67LJEIuFMVi6Xa7GnMlWRxsbGhQ8+8cQTFrSIpaIfGC7EsEkCw4YxbBJd3KqiX9fPXNlkJVqeuiDuXxU038li30IMsko/LrJ58+bMNxWK+VHsaRJAWeFz/pq/csqqwAVxs+KTVfpjkV27dvX09PT09GzatKmtrc3r9YrHQ6HQ2NjY4OBgZ2cnpj8BMINPn5G/ssnKOGXWivhIXb1ll2bp+GSVvoo0NjYeOHDA7/f39fW9+OKL4rjE7XavW7furrvuOnDgQNqTXQCwEJ8+I39lk9WvxuPkJSJyRUa3b7e+hBCnrNJXkd7eXnGosXLlShxzABSCT5+Rv/LIanBgaN6bvBwr7wVxs+KTVfoq8vTTT69cuTIcDu/evXvhYQdm0AIwj0+fkb/yyOr32nqx4Z3Ze3dhi4hkwCer9FWkr6/vZz/7WSgUmpycXHgsgioCYB6fPiN/ZZDVsz98NXHtFrJ0yqy0+GSVvoqIud99Pt/u3btxRgugEHz6jPyVelbTgem583eqLzu1v67+K/Z9F5+sMt0v0trailmzAArEp8/IX6lnZVwQ91t/Z2MJIU5ZZZlHCwAKxKfPyF9JZ/Xa/oMWLoibFZ+sUEUA7MVn7lX+Sjqrd2eXiY2q0MFNWwpdEDcrPlmhigDYi0+fkb/Szerll/dG3e1i+441igPfyCcrVBEAe/HpM/JXullZviBuVnyyQhUBsBefPiN/JZqVviCuTVNmpcUnK1QRAHvx6TPyV4pZjR7y6QviVs+PWLUgblZ8skIVAbAXnz4jf6WY1ZvvRsSCuK7IqIUL4mbFJytUEQB78ekz8ldyWRkXxF2pTjn51XyyQhUBsBefPiN/JZfVHxLJURD3qf0OXN1rxCcrVBEAe/HpM/JXWln1vbQ37lpNRJIW6b63yeFv55MVqgiAvfj0GfkroaymA9Mz3vvEtmf2DccG1XV8skIVAbAXnz4jfyWU1Z5942JQXY5NP2bb9O8Z8MkKVQTAXnz6jPyVSlYpC+IWpQ18skIVAbAXnz4jf6WS1evvJ+c4sW9B3Kz4ZIUqAmAvPn1G/koiq8GBIX3KrDtXx4vVDD5ZoYoA2ItPn5G/ksjKuCCuM1NmpcUnK1QRAHvx6TPyxz+rZ3/4qpgyy+4FcbPikxWqCIC9+PQZ+WOelXFB3JrQPuev7jXikxWqCIC9+PQZ+eOcVWju493/9r64uleJvu/klFlp8ckKVQTAXnz6jPxxzuqnvc9EPp2c4+Szbn9xG0OcspIt/8RgMDgyMhKLxdrb25ubmy3/fIDSwqfPyB/brJ5/6DuzbfeIbff8yFe2fam47SFOWVl8LBKNRgcHB9vb2zs7O4eHh0OhkLWfD1By+PQZ+eOZ1b9++9nT6lVR7zoikrTIljvri90iIk5ZWXws8v777zc2NjY2NhJRc3PzxMRER0fHYi+eePmNmd8HrG1A4WRJSmhasVtxEVmSlspLz8T/u9gNScUwK0miS+VLuWUlSUREzKKiy5RLiYhhVqyCCsci0dcmQ4/vEH9+xjXx6WvXFrdJAp9jEYurSCgUUlVV/zMWixGR3+/3+/1EtPqGGzxXXqk/O/JvvwscPGJtAwAArPXxw3eJq3uVxIdb77nJ7XYXu0VJkiRpTpVcSVr0xBVG1wEAFpWod4du2Cq2O1YEjP3gosvwy+4k60fXF9LPcQWDwUgkoj/e9MDnqlZd60ADcsLwLA3OaJnHMytJIonYZcX2jBbZdvYvHIuEon/+c+xMJP7JfCIWjUezvsVz2wNiw6tN3H7LTcZfsKJTVTWRSCQSCQe+y+v1LvaUxVWkubl5aGhIbAeDwfb29kwv3v6lGn7D76qqihNxfKiq6nK5GF6qwDMrRVFY/VMnIlmWZVnmlpU4M8MwKyKy5JdxMjQ9Mx86Fgp8OB+amf9oMjSdfOLC2EvqD2BN1bKaKu9nvJ+urVrW5K2bDcR/OZW80PSvbq0qvEkWKvqgus7iKlJbW1tbWzswMEBEXq8XV/oC6LRz54o+EFoq8shqZj40Mz/nmzn2STw2M//R4Zlji3968kinSnE1eetEzaip8q6tWWV8VSwhveKPibP+K+Xf3Hj9nawqrvEareLuV9af0erq6goGg0RUW1tr+YcDlBw+19LwZz6r+XhsMhSYDE1H4tEPQn/KVDMM1tas+oz3027F1eSta/LWVylqhhe/+utJTWomIjk+ve2BTKdVioLPfmXLuAjqB4COT5+RvwxZHZ45NjMfCs7PfRD602Roet7EkMbamlU1VddcoaitNavEqSrzLTkWjB2NJE+l3Fbt81y5yfx7ncFnv3JidB2gkvHpM/KnZ+U/fXIyNC1qhjhVlfW9Td66mqprrq3yrvLW11R5m7wFTZX46lhyYMY9P7LhwbsL+Sib8NmvUEUA7MWnz8jWRcPg0Y8n57LfjJwyDF5gzUixzxf75OwyIpK0yOb1Fn6wlfjsV6giAPbi02dkIodh8PMyD4NbKxiWfhfwiO0bEv31TQ/b912F4LNfoYoA2ItPn7EoHBgGt9aet6NEChG55t+6r4vjuSyBz36FKgJgLz59RmcUOAx+7VXVyy9f5syddAsNT8hzMQ8RSVrktprxKs/2ojTDDD77FaoIgL349BntYKwZlgyDi7sOi5KVf076zbHkT+Ky0I/a7/ufDjcgJ3z2K1QRAHvx6TMWbtG7wReX0zB4EbOKJaTdhxSxXRXZf0/rMkXlMutiWnz2K1QRAHvx6TPmyvlh8CJmNXBIjsWJiOT49A2JV5paX3Ty2/PAZ79CFQGwF58+Y2b6MPjJ+ZDJmkFWD4MXK6uRSfn4bHJ+3GtPfqf9f/yNk9+eHz77FaoIgL349BlTvPOno7//6ERoPuzM3eBmFCWrYFg6MJ78JfTOvdRyfW1NA9ebRAz47FeoIgD2YtJntHwY3A7OZxVLSLtHk8Mh6n+PLz/d17zp/znz1QVisl8RqgiA3YrSZ7R7GNwmzmc1cEgOzxMRSVrk2sATTW2bqjxF+A/PA45FACqFA31G4zD4ZOiEmZpRpbiuv6Zh9TUrliluu+8GN8/h/vXQYUUfDll+4jGv67/XfukxB77XEjgWAagUlvcZrRoGZ7hKlZP969Gppe9MJUtI9cl/rIq+tfbe79n6jdbCsQhApSi8z3h45pg+gwiTYXCbONa/Doalnx++RGx7QgNXhV6oXrm+fjXf+U4WwrEIQKXItc9YEsPgNnGmfx1LSP/81oUR9WtOfY+ISuhcloBjEYBKkbnPWKLD4DZxoH8dS0j/PJK8wVDSIp/2f11KRJq/8LC3psSW98axCEClMPYZZ//rdB7D4I5Nil50dvevRQkJhiUikpec+fTxrVIiorjczZ//uuXfZTcciwCUP+Mw+EdnTo+dnDDzriJOil50dvevBw4lSwgRVc+8oMSOE1HzrV9nPmVWWjgWAShDlTMMbhNb+9fG63pbaO9f5p7TiK6ua27+PNN1qDLDsQhAyStwGLzJW1fhNWMh+/rXQ4cV/breOxsDwZ/3ahoR0U13lNigug7HIgAlppBh8NXLGuo81fVV1Q60s6TZ1L82lpBbGrTE+HdECam/8e6SmDIrLRyLALCW393giw2Dy7Isy3IsFrOzyeXAjv51SglpUYdGjr8t/iy5q3uNcCwCwEged4PrNcP8MHjR+4z8Wd6/Tikhd6+e2/9Pz4g/m7/wcKlMmZUWjkUAisnJYXA+fUb+rM0qpYR0rY0ffuPH86FpIqry1pXi1b1GfFPWLSgAAA70SURBVPYrVBEof8UdBufTZ+TPqqxiCenVw/LR6YtKyHx4enJsUDxSolf3GvHZr1BFoHyI3tkHHwU+OBXIbxjcjhlE+PQZ+bMkK+OthXS+hBDRxFs/jsciRFSzcn1T66aCG1tkfPYrVBEobdYOg9uBT5+Rv8KzCoalf35LEROcENFdaxIdTQkimpl6Wz8QKelBdR2f/crRKuJyuS76bllOeYQDSZK4tUpc4cOtVVSMrObj0YnZqcnQ9PSfT83Mh8zcDV6lXn791StWX7NCFI/m6oYqpQhJSpJUdfnlzn9vBrIs04J/lUwoipLHu96dWvKvo2f1EvLVddL6JoVIIaL//NUzRCRJ9Jm27rrG9lw/mWdWkiRJkpRfVhZytIpEoxeNYaqqmvIIB6qqcrsiU1VVWZYrM6s8hsHbljdfc9lViw6DJyiacDpJnlf6il9GbvuVaFUikcj1jSOTsr58uqrQ1z4fr/Vo4j9u0jf48fQEEcmK+7MbHs3jP5lnVqqqJhKJPLLK77sWewpntICRwofBG5bVKorCauUl/fw1ZJVfVilj6bUe7f61iVqPJv6MxyKHX09e3dvUtqnUB9V1fPYrVBEomkLuBi+hhTT4nL/mL4+sgmFp96gi1k4nolqP9rWOhCpr+gv0QfUqb115jIgIfPYrVBFwCP9hcJvwuZaGv1yzGp1aemDiEn0g5LZVic7mi07vzIenJ377I7G97svfsayhDPDZr1BFwBbO3A1eEvj0Gfkzn1XKWSxVoQfWJpprU0cIDu37rtioWbm+dKfMSovPfoUqAtbApOiL4dNn5M9kVhNB+d8Oy/ohSK1H616X8Li0lJfNTL09c37KrHX3ldWBCHHar1BFIB9iGHwuHjk6cxyTomfGp8/IX9asUg5BKN1ZLN2hf08eiJT6lFlp8dmvUEUguwoZBrcJnz4jf5mzShkF8VTRA2vjjctSD0GEw288I6bMUtSSXBA3Kz77FaoIpKrYYXCb8Okz8rdYVsGwNDwh6ysVEtFtqxIdqzTjtVhG8Vjkwp3qdz5WNlf3GvHZr1BFKl0hw+DeKk9DVXXZDIPbhE+fkb+FWcUS0sgx6TfHLvxSZT4EEQ6/8Yy4uvfquuYymDIrLT77FapIxbFwGFxV1eiZM0XfiZnj02fkLyWrkUn5N8cujKJTtkMQITQzoR+IlO6CuFnx2a9QRcqc3ZOi42cxKz59xpKQSCRkWfadUH4zeYl+LyERrazWvrI2zYVYCx1+48Kd6mV2da8Rn/0KVaSsGIfBJ0N/MlMzMAxuNz59xpIwOrU0pX54quie5jT3gqQ16RsUV/eW66C6js9+hSpSwjAMXhL49BmZS1s/bms6297wF5OfEI9FJv7jx2K7qW1T+V3da8Rnv0IVKRm4G7xE8ekz8hRLSKNT0jtTsrF+qArd0pB9CCTFxFvlsyBuVnz2K1QRpvSaEYlH/X/+8A8fn5iP/VfWd1XI3eClhU+fkZtwVBo5Jr8XlIzj5/nVD7p4yqxyvbrXiM9+hSrCRXHXBgf78Okz8jE6tXQiuMR4/wedP3+1pu5slSolErmVEDIMqtesXF+/+m5rGsoYn/0KVaQ4cDd45eDTZyy6YFiaCErvTF108S4R1Xq0toZzYvxDrAeVq5mptwNHXxPb5TT9ewZ89itUESfow+Dma4ZxGLzOW/PFptZQKORAU8FyfPqMxRKOSuPB1JEP4ZYGrbXhwopSulyzEgviElFT2yZvTXMBjS0ZfPYrVBHriSENcemUJcPgGdaqBP749BkdJorH0WkpGJZSnqr1aDfWae0NqYMf+WU18daPxIK4iuqukAMR4rRfoYpYAJOiQwZ8+ozO8M9J/lnp+Gya4iFGzptrtYUHH0IeWcVjkYnfJq/ubf7C18t+UF3HZ79CFckZhsEhJ3z6jPYJR6XJ2Uv8c0uOz0opYx5EpCr02VqtsVrLeudgHlkZF8Rt/vzDubW7lPHZr1BFssAwOBSIT5/RWqJyzISXfDAnLRzwoFyKhy7XrMp4Qdys+OxXqCIXmZkPzYUihwLjuBscrMKnz1g4/5wUDEvBsHQynL5yEFGtR1tZrWU4bZVBrlnpC+LW33h3GU+ZlRaf/aqiq4jlw+AAC/HpM+bBPyeF5i8JR88Fw1LK7R1Gnir6zDKtxnOuqfqsmQkTF5NTVoH3X9MXxK2cQXUdn/2qsqoIhsHBeXz6jFkFw1IsQf5ZKRSVTs/TwuFxI1E5PFXamlqtkMphlFNWh19PXt1blgviZsVnvyrnKoJhcOCAT5/RKJaQZma02bA2G5bN1AxhZbVW69FqPVTrsaxyGJnPquwXxM2Kz35VPlXEqmFwVVVjsZj97YVKUdw+YywhBcNERP5ZiShZKs6fmxKD3pl+BETZ8LiW1HrO5jHOkSuTWc2Hp8t+QdyscCxSqALvBscwODjG7j6jXifEAAadLxUfRxcdAF/MympNVcjrsvFoIzOTWelX99asXF+uC+JmhWOR3GBSdChdefcZxSiF4D8/sh2LS6EoEdGZuKlzUIup9Whu1yXXekiW/lLr0bwucr5mLGQmq5mpt/UDkco8lyWU/7HI8PBwZ2dn3m/HMDiUB/+c+KHX5Eu0xFmJDPVA0KuCcDKc5q69Qqys1ohInIyq9ZCqaLUeEvOOuN1uIopEzlj5fYUx07+eeOvCOlSVdnWvUTkfi/j9/tHR0ZyGFgoZBkfNgMUY+/Jm6GeEMn9m2sfNHRkoObTGHFWh5R6NiLwuUpULpUKVyYFhDMtl7V9XzoK4WZXzsYjf73e5XGaqyI4f/WQ+Kie0C//QZVpeR8sXXrInS7JLvpSIqpTLLpeXyvJSIqJ5onk6HNT+S4sTzVj3X8CSJJFWej8KOTlzruqTs8us+KQ4UWmfvZSXnPFeciq5LcWukMJi+2rppNi4QvrYJZ1Ovjpm+F+i+AzFiSJEs9m+RVUUIorFLT32sZl+LqvsF8TNqpyPRTo7O/1+/8jIiP6Iz+cbHx8nog1fvL265lPiQUmSzsbqNdd6M2d2NSIxTDhPRInzl5YAsCRpETU6nvKgEp+WtIjxEfXMe5ckLjyixsalRIQy+njBRgWSJHJ56r606Uknv1SsepLf2if2kSSJiDRH+pfiu9JyIpSGhgaPx0NEsrI0Gr1wDvjskioHvh3KVdofa/PvVeLZr+sjoss/+W2GZ13zb+XXAMibrLrX3fO48ZfEAaqqSpLk8JdmJctyIuFQn9rlci3ajLw/VD/C0G3fvj3tK71er9frJaJgMKif6dLOnbv/Fvnjj393+aWX5d0GOzj5f4xJsiwrqhKd57UHE5Esyy553q0wathSWVmqKNFo1ktcVaImcx+Z+WVfM/MRkiRdIsl/SfA6d+RyVRGRiawclbV/7a1pVlS3wzd1KYqiaRqfO8nEWSxVVROJhDO/V7ZUkdbW1tbW1rzfLi1Z8sV2juv3MbzrUFVVl8uFrMxQVVVRlEgky6khh8myLMsyt6zOX6PFLisi4taT48Y4EFL0cZH8rzcHADP0a2kgK2RlXjlfo0VEjY2NjY2NdnwyQMnhcy0Nf8jKPD5Z4VgEwF58+oz8ISvz+GSFKgJgL2Ofsbgt4Q9ZmccnK1QRAHvx6TPyh6zM45MVqgiAvfj0GflDVubxyQpVBMBefPqM/CEr8/hkteScg3Us5Xr5gYGBLVu2OPbtpWtqaurQoUPIyoypqampqakNGzYUuyEl4ODBg0SErMxAVoqiLDYJiqPTwqhq6hx5Cx+BhRRFIWRljqIosiwjKzPE/X3IygxklQHOaAEAQP4cPaMFAABlBsciAACQP6enyzeupBsMBkdGRmKxWHt7e3Nzs/FlGZ6qHCmzJtfW1hoXIR4aGtLnZ+zs7KytrXW6fcxkCAS7U4qJiYnR0VFasFMR9isiyrbDYHdK4VwVSVlJNxqNDg4OdnV1qao6ODhYU1MjZo/P/FRF0ddlIaLR0dGUEN55552vfS05Lbn+skq2WCDYnVKEQqGhoaFt27aJQHw+n3FybuxXmXcY7E4LXfIP//APznzT2NjYkiVLzpw5c9NNNxHRe++9J8vy5z73uSuuuOKTTz45ffp0fX29eGWGpyrKZZdddtVVV1111VVnzpw5duzYxo0b9aei0ejMzMyGDRsuvfTS6urqpUuXFrGdHGQIBLtTirGxsWuuueamm2667LLLzpw589FHH61atUo8hf2Ksu0w2J0Wcu5YJGUl3VAoZLxszngrSYanKtPw8HBXV5fxkZmZmZmZmRdeeIGIVFXdtGlThjVkKkGGQLA7pejo6BAb0WjU7/evWbNGfwr7FWXbYbA7LcRrGWFYyO/364tF6jweT1dXlzgnOzAw4PP59J+GyoRAcjUxMTE8PNzc3Gw8nYUYIQ/WVxHzK+mCLkNoo6OjC/8lG+tKe3v7yMhI5fxrT5tVJQeSwWL71dDQUDQa3bZtW0rvBDFCHqyvIiZX0m1ubh4aGhLbwWCwvb2diPx+f2NjY9qnyttioYnz1MbrZEREw8PDRCSurhGPONbUokubVdpAKnZ30qXNyufzRaPRlAl1sF/pMvw0LfZshXNudJ2ITp8+HQgExOj6FVdc8eGHHx4+fPjo0aNut/sLX/gCEX3/+9+/44470j5Vmd577714PH7jjTfqj4iIampq3nzzzePHjx89evTDDz+85557KnMgVJc2EOxOaR09enR0dPTgeWfPnm1sbMR+pcvw07TYsxWuyPeuB4NBIkp7TXqGp0Dw+/2qqiIiXYZAsDuZh/2Ksu0w2J2MMAMKAADkDzOgAABA/lBFAAAgf6giAAXx+/0+n0//c3h42O/3F7E9AA5DFQEoSDgcbmtrE4VkYGCgp6enMqefgoqF0XWAQvX29g4PD7/++usrVqzYs2dPyiy5AOUNxyIAhdq1a5fX612xYsWjjz6KEgKVBlUEwAKbN2+ORCLixjSAioIzWgCFCoVC11133Te/+c3h4eGxsbFiNwfAUTgWAShUd3f3o48+umvXLiLq7e0tdnMAHIVjEYCCPP300319fe+8847X6/X5fG1tbWNjY2YmJAUoD6giAACQP5zRAgCA/KGKAABA/lBFAAAgf6giAACQP1QRAADI3/8H96X1ymUjyxQAAAAASUVORK5CYII="}}},{"cell_type":"code","source":"def leaky_relu(x):\n    return tf.nn.relu(x) - 0.01 * tf.nn.relu(-x)","metadata":{"_cell_guid":"d697dbb2-54e2-49aa-93dd-0c7c0cd12996","_uuid":"51764360ba91ecdff96c6c89bb57142c6f19bac7","execution":{"iopub.status.busy":"2022-08-01T10:35:15.464447Z","iopub.execute_input":"2022-08-01T10:35:15.464785Z","iopub.status.idle":"2022-08-01T10:35:15.471050Z","shell.execute_reply.started":"2022-08-01T10:35:15.464699Z","shell.execute_reply":"2022-08-01T10:35:15.469621Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Model\nregressor = tf.contrib.learn.DNNRegressor(feature_columns=feature_cols, \n                                          activation_fn = leaky_relu, hidden_units=[200, 100, 50, 25, 12])\n    \n# Deep Neural Network Regressor with the training set which contain the data split by train test split\nregressor.fit(input_fn=lambda: input_fn(training_set), steps=2000)\n\n# Evaluation on the test set created by train_test_split\nev = regressor.evaluate(input_fn=lambda: input_fn(testing_set), steps=1)","metadata":{"_cell_guid":"49fe448a-883e-4693-b5c0-50417488336f","_uuid":"940beb55da2e02f704c47a21d7ca4490a8852419","execution":{"iopub.status.busy":"2022-08-01T10:35:25.814496Z","iopub.execute_input":"2022-08-01T10:35:25.815005Z","iopub.status.idle":"2022-08-01T10:35:32.894741Z","shell.execute_reply.started":"2022-08-01T10:35:25.814941Z","shell.execute_reply":"2022-08-01T10:35:32.893824Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the score on the testing set\n# 0.002X in average\nloss_score2 = ev[\"loss\"]\nprint(\"Final Loss on the testing set with Leaky Relu: {0:f}\".format(loss_score2))","metadata":{"_cell_guid":"c3885a56-a0a1-434e-b6f4-d33440d3b721","_uuid":"ff854c7ca8e0c03e0cf86edd3ad566fc79ebc05f","execution":{"iopub.status.busy":"2022-08-01T10:35:49.723105Z","iopub.execute_input":"2022-08-01T10:35:49.723596Z","iopub.status.idle":"2022-08-01T10:35:49.730485Z","shell.execute_reply.started":"2022-08-01T10:35:49.723483Z","shell.execute_reply":"2022-08-01T10:35:49.729365Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Predictions\ny_predict = regressor.predict(input_fn=lambda: input_fn(test, pred = True))\nto_submit(y_predict, \"Leaky_relu\")","metadata":{"_cell_guid":"4df61196-e5e6-400b-90b4-f1c9736f92ab","_uuid":"c96f4a907c614cd2bf2cfd649c8869e27aeac512","execution":{"iopub.status.busy":"2022-08-01T10:35:55.753399Z","iopub.execute_input":"2022-08-01T10:35:55.753768Z","iopub.status.idle":"2022-08-01T10:35:56.227462Z","shell.execute_reply.started":"2022-08-01T10:35:55.753670Z","shell.execute_reply":"2022-08-01T10:35:56.226575Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Model\nregressor = tf.contrib.learn.DNNRegressor(feature_columns=feature_cols, \n                                          activation_fn = tf.nn.elu, hidden_units=[200, 100, 50, 25, 12])\n    \n# Deep Neural Network Regressor with the training set which contain the data split by train test split\nregressor.fit(input_fn=lambda: input_fn(training_set), steps=2000)\n\n# Evaluation on the test set created by train_test_split\nev = regressor.evaluate(input_fn=lambda: input_fn(testing_set), steps=1)\n\nloss_score3 = ev[\"loss\"]","metadata":{"_cell_guid":"d33bd73c-5ebb-41bf-911c-bd8e6a4e6817","_uuid":"83507853824f10fa77ce98e9f78d70dcbd857149","execution":{"iopub.status.busy":"2022-08-01T10:36:04.344484Z","iopub.execute_input":"2022-08-01T10:36:04.344885Z","iopub.status.idle":"2022-08-01T10:36:10.415636Z","shell.execute_reply.started":"2022-08-01T10:36:04.344808Z","shell.execute_reply":"2022-08-01T10:36:10.414617Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Final Loss on the testing set with Elu: {0:f}\".format(loss_score3))","metadata":{"_cell_guid":"97a3110c-cd2f-47b1-a17e-2256ec224f55","_uuid":"7dd50347b565abc2d23486fda6488385116fe637","execution":{"iopub.status.busy":"2022-08-01T10:36:13.694841Z","iopub.execute_input":"2022-08-01T10:36:13.695185Z","iopub.status.idle":"2022-08-01T10:36:13.700620Z","shell.execute_reply.started":"2022-08-01T10:36:13.695124Z","shell.execute_reply":"2022-08-01T10:36:13.699361Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Predictions\ny_predict = regressor.predict(input_fn=lambda: input_fn(test, pred = True))\nto_submit(y_predict, \"Elu\")","metadata":{"_cell_guid":"ca9e2de3-12e8-4881-92ba-f51fe185c958","_uuid":"42ba57572d89bf477957f9d8e677759b77620f40","execution":{"iopub.status.busy":"2022-08-01T10:36:17.221516Z","iopub.execute_input":"2022-08-01T10:36:17.221917Z","iopub.status.idle":"2022-08-01T10:36:17.685821Z","shell.execute_reply.started":"2022-08-01T10:36:17.221847Z","shell.execute_reply":"2022-08-01T10:36:17.684987Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So we have 3 submissions with 3 differents activation functions. But we built ours models just with the continuous features. If you want to comapare the performance download the 3 submissions and submit to the leaderboard. \nNow we see how do build another model by adding a categorical features.","metadata":{"_cell_guid":"e49c1375-ff65-49e2-a594-b97c1cfffea1","_uuid":"c131e5e4af5379787e492e1eb2a1855c27e93d30"}},{"cell_type":"markdown","source":"# <center> VII. Deep Neural Network for **continuous and categorical** features","metadata":{"_cell_guid":"f8b3a45c-b978-4861-a44f-1236be9d4d2e","_uuid":"0ca9015e24c91d4ea1503baf7002a2ecdc888d1b"}},{"cell_type":"markdown","source":"For this part I repeat the same functions that you can find previously by adding a categorical features.\n","metadata":{"_cell_guid":"c607487a-83a1-4ba9-bd8f-ec9392fb1326","_uuid":"173e1941a581e7a12d3054238af555d1fec89183"}},{"cell_type":"code","source":"# Import and split\ntrain = pd.read_csv('../input/train.csv')\ntrain.drop('Id',axis = 1, inplace = True)\ntrain_numerical = train.select_dtypes(exclude=['object'])\ntrain_numerical.fillna(0,inplace = True)\ntrain_categoric = train.select_dtypes(include=['object'])\ntrain_categoric.fillna('NONE',inplace = True)\ntrain = train_numerical.merge(train_categoric, left_index = True, right_index = True) \n\ntest = pd.read_csv('../input/test.csv')\nID = test.Id\ntest.drop('Id',axis = 1, inplace = True)\ntest_numerical = test.select_dtypes(exclude=['object'])\ntest_numerical.fillna(0,inplace = True)\ntest_categoric = test.select_dtypes(include=['object'])\ntest_categoric.fillna('NONE',inplace = True)\ntest = test_numerical.merge(test_categoric, left_index = True, right_index = True) ","metadata":{"_cell_guid":"5ae69e94-625c-421b-b3b5-325f80aa0164","_uuid":"b536485388dc43bc383179d02e6c89b054f37159","execution":{"iopub.status.busy":"2022-08-01T10:37:25.696343Z","iopub.execute_input":"2022-08-01T10:37:25.696698Z","iopub.status.idle":"2022-08-01T10:37:26.402168Z","shell.execute_reply.started":"2022-08-01T10:37:25.696640Z","shell.execute_reply":"2022-08-01T10:37:26.401356Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-01T10:37:37.129435Z","iopub.execute_input":"2022-08-01T10:37:37.129832Z","iopub.status.idle":"2022-08-01T10:37:37.169650Z","shell.execute_reply.started":"2022-08-01T10:37:37.129765Z","shell.execute_reply":"2022-08-01T10:37:37.168819Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-01T10:37:57.418030Z","iopub.execute_input":"2022-08-01T10:37:57.418392Z","iopub.status.idle":"2022-08-01T10:37:57.460935Z","shell.execute_reply.started":"2022-08-01T10:37:57.418332Z","shell.execute_reply":"2022-08-01T10:37:57.459889Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Removie the outliers\nfrom sklearn.ensemble import IsolationForest\n\nclf = IsolationForest(max_samples = 100, random_state = 42)\nclf.fit(train_numerical)\ny_noano = clf.predict(train_numerical)\ny_noano = pd.DataFrame(y_noano, columns = ['Top'])\ny_noano[y_noano['Top'] == 1].index.values\n\ntrain_numerical = train_numerical.iloc[y_noano[y_noano['Top'] == 1].index.values]\ntrain_numerical.reset_index(drop = True, inplace = True)\n\ntrain_categoric = train_categoric.iloc[y_noano[y_noano['Top'] == 1].index.values]\ntrain_categoric.reset_index(drop = True, inplace = True)\n\ntrain = train.iloc[y_noano[y_noano['Top'] == 1].index.values]\ntrain.reset_index(drop = True, inplace = True)","metadata":{"_cell_guid":"f67a2138-cdbc-4282-bd1d-55cc70c5c466","_uuid":"8b5200d119fb5f1e836fb746e63bf26a5cb53aea","execution":{"iopub.status.busy":"2022-08-01T10:38:54.991492Z","iopub.execute_input":"2022-08-01T10:38:54.991896Z","iopub.status.idle":"2022-08-01T10:38:55.337665Z","shell.execute_reply.started":"2022-08-01T10:38:54.991828Z","shell.execute_reply":"2022-08-01T10:38:55.336756Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-01T10:39:05.859144Z","iopub.execute_input":"2022-08-01T10:39:05.859468Z","iopub.status.idle":"2022-08-01T10:39:05.867699Z","shell.execute_reply.started":"2022-08-01T10:39:05.859410Z","shell.execute_reply":"2022-08-01T10:39:05.866592Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"col_train_num = list(train_numerical.columns)\ncol_train_num_bis = list(train_numerical.columns)\n\ncol_train_cat = list(train_categoric.columns)\n\ncol_train_num_bis.remove('SalePrice')\n\nmat_train = np.matrix(train_numerical)\nmat_test  = np.matrix(test_numerical)\nmat_new = np.matrix(train_numerical.drop('SalePrice',axis = 1))\nmat_y = np.array(train.SalePrice)\n\nprepro_y = MinMaxScaler()\nprepro_y.fit(mat_y.reshape(1314,1))\n\nprepro = MinMaxScaler()\nprepro.fit(mat_train)\n\nprepro_test = MinMaxScaler()\nprepro_test.fit(mat_new)\n\ntrain_num_scale = pd.DataFrame(prepro.transform(mat_train),columns = col_train)\ntest_num_scale  = pd.DataFrame(prepro_test.transform(mat_test),columns = col_train_bis)","metadata":{"_cell_guid":"0639e24d-202b-461c-b708-0c7d5bc3aa18","_uuid":"b4b2a5534771950433404e618d6b002a913606a2","execution":{"iopub.status.busy":"2022-08-01T10:39:24.651872Z","iopub.execute_input":"2022-08-01T10:39:24.652304Z","iopub.status.idle":"2022-08-01T10:39:24.690734Z","shell.execute_reply.started":"2022-08-01T10:39:24.652221Z","shell.execute_reply":"2022-08-01T10:39:24.689864Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(col_train_num[:5])\nprint(\"-----------------------------------------------------------------------------------------------\")\nprint(col_train_num_bis[:5])\nprint(\"-----------------------------------------------------------------------------------------------\")\nprint(col_train_cat[:5])\nprint(\"-----------------------------------------------------------------------------------------------\")\nprint(mat_new[:5])\nprint(\"-----------------------------------------------------------------------------------------------\")\nprint(mat_train[:5])\nprint(\"-----------------------------------------------------------------------------------------------\")\nprint(mat_test[:5])\nprint(\"-----------------------------------------------------------------------------------------------\")\nprint(prepro)\nprint(\"-----------------------------------------------------------------------------------------------\")\nprint(prepro_test)\nprint(\"-----------------------------------------------------------------------------------------------\")\nprint(prepro_y)\nprint(\"-----------------------------------------------------------------------------------------------\")\nprint(train_num_scale.head())\nprint(\"-----------------------------------------------------------------------------------------------\")\nprint(test_num_scale.head())","metadata":{"execution":{"iopub.status.busy":"2022-08-01T10:46:47.528099Z","iopub.execute_input":"2022-08-01T10:46:47.528446Z","iopub.status.idle":"2022-08-01T10:46:47.605920Z","shell.execute_reply.started":"2022-08-01T10:46:47.528384Z","shell.execute_reply":"2022-08-01T10:46:47.605174Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train[col_train_num] = pd.DataFrame(prepro.transform(mat_train),columns = col_train_num)\ntest[col_train_num_bis]  = test_num_scale","metadata":{"_cell_guid":"8ae00e60-3fe9-4297-a36e-91be9af657e0","_uuid":"666dc634a6f1b56d4e5492644ecb49d9fc444fc6","execution":{"iopub.status.busy":"2022-08-01T10:46:56.358457Z","iopub.execute_input":"2022-08-01T10:46:56.358811Z","iopub.status.idle":"2022-08-01T10:46:56.522555Z","shell.execute_reply.started":"2022-08-01T10:46:56.358723Z","shell.execute_reply":"2022-08-01T10:46:56.521577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The principal changements are here with the lines beginning by for categorical_features... It is possible to use other function to prepare your categorical data. ","metadata":{"_cell_guid":"06a63dfe-4b39-4281-a240-7620595dcb87","_uuid":"043062c7d11e3503cf2d60da1e3eb46639228a36"}},{"cell_type":"code","source":"FEATURES","metadata":{"execution":{"iopub.status.busy":"2022-08-01T10:48:01.927151Z","iopub.execute_input":"2022-08-01T10:48:01.927478Z","iopub.status.idle":"2022-08-01T10:48:01.933882Z","shell.execute_reply.started":"2022-08-01T10:48:01.927420Z","shell.execute_reply":"2022-08-01T10:48:01.932872Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# List of features\nCOLUMNS = col_train_num\nFEATURES = col_train_num_bis\nLABEL = \"SalePrice\"\n\nFEATURES_CAT = col_train_cat\n\nengineered_features = []\n\nfor continuous_feature in FEATURES:\n    engineered_features.append(\n        tf.contrib.layers.real_valued_column(continuous_feature))\n\nfor categorical_feature in FEATURES_CAT:\n    sparse_column = tf.contrib.layers.sparse_column_with_hash_bucket(\n        categorical_feature, hash_bucket_size=1000)\n\n    engineered_features.append(tf.contrib.layers.embedding_column(sparse_id_column=sparse_column, dimension=16,combiner=\"sum\"))\n                                 \n# Training set and Prediction set with the features to predict\ntraining_set = train[FEATURES + FEATURES_CAT]\nprediction_set = train.SalePrice\n\n# Train and Test \nx_train, x_test, y_train, y_test = train_test_split(training_set[FEATURES + FEATURES_CAT] ,\n                                                    prediction_set, test_size=0.33, random_state=42)\ny_train = pd.DataFrame(y_train, columns = [LABEL])\ntraining_set = pd.DataFrame(x_train, columns = FEATURES + FEATURES_CAT).merge(y_train, left_index = True, right_index = True)\n\n# Training for submission\ntraining_sub = training_set[FEATURES + FEATURES_CAT]\ntesting_sub = test[FEATURES + FEATURES_CAT]","metadata":{"_cell_guid":"268dc739-d13f-42e5-8eb0-f0f16e75b41b","_uuid":"dae02f368a2cc5b5ca248e55e5b63e1e3477993e","execution":{"iopub.status.busy":"2022-08-01T10:49:05.148168Z","iopub.execute_input":"2022-08-01T10:49:05.148513Z","iopub.status.idle":"2022-08-01T10:49:05.206450Z","shell.execute_reply.started":"2022-08-01T10:49:05.148439Z","shell.execute_reply":"2022-08-01T10:49:05.205422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"engineered_features","metadata":{"execution":{"iopub.status.busy":"2022-08-01T10:49:23.900161Z","iopub.execute_input":"2022-08-01T10:49:23.900500Z","iopub.status.idle":"2022-08-01T10:49:23.909525Z","shell.execute_reply.started":"2022-08-01T10:49:23.900440Z","shell.execute_reply":"2022-08-01T10:49:23.908442Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"training_sub.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-01T10:50:05.270679Z","iopub.execute_input":"2022-08-01T10:50:05.271035Z","iopub.status.idle":"2022-08-01T10:50:05.307639Z","shell.execute_reply.started":"2022-08-01T10:50:05.270976Z","shell.execute_reply":"2022-08-01T10:50:05.306263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"testing_sub.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-01T10:50:35.564691Z","iopub.execute_input":"2022-08-01T10:50:35.565069Z","iopub.status.idle":"2022-08-01T10:50:35.601311Z","shell.execute_reply.started":"2022-08-01T10:50:35.564989Z","shell.execute_reply":"2022-08-01T10:50:35.600191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Same thing but for the test set\ny_test = pd.DataFrame(y_test, columns = [LABEL])\ntesting_set = pd.DataFrame(x_test, columns = FEATURES + FEATURES_CAT).merge(y_test, left_index = True, right_index = True)","metadata":{"_cell_guid":"b7a855c8-82d4-4935-83c0-b05f6d917a7e","_uuid":"9be1327ff12a8c1e0a3ef6920cc482fea1a6b346","execution":{"iopub.status.busy":"2022-08-01T10:50:42.467548Z","iopub.execute_input":"2022-08-01T10:50:42.467881Z","iopub.status.idle":"2022-08-01T10:50:42.477298Z","shell.execute_reply.started":"2022-08-01T10:50:42.467822Z","shell.execute_reply":"2022-08-01T10:50:42.476308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**What is SparseTensor?**\n\nA sparse tensor is a dataset in which most of the entries are zero, one such example would be a large diagonal matrix. (which has many zero elements). It does not store the whole values of the tensor object but stores the non-zero values and the corresponding coordinates of them.\n\n**What is a TF SparseTensor?**\n\nA nonzero value in the context of a tf. sparse. SparseTensor is a value that's not explicitly encoded. It is possible to explicitly include zero values in the values of a COO sparse matrix, but these \"explicit zeros\" are generally not included when referring to nonzero values in a sparse tensor.\n\nhttps://www.tensorflow.org/guide/sparse_tensor\n\nhttps://www.tensorflow.org/api_docs/python/tf/sparse/SparseTensor","metadata":{}},{"cell_type":"code","source":"training_set[FEATURES_CAT] = training_set[FEATURES_CAT].applymap(str)\ntesting_set[FEATURES_CAT] = testing_set[FEATURES_CAT].applymap(str)\n\ndef input_fn_new(data_set, training = True):\n    continuous_cols = {k: tf.constant(data_set[k].values) for k in FEATURES}\n    \n    categorical_cols = {k: tf.SparseTensor(\n        indices=[[i, 0] for i in range(data_set[k].size)], values = data_set[k].values, dense_shape = [data_set[k].size, 1]) for k in FEATURES_CAT}\n\n    # Merges the two dictionaries into one.\n    feature_cols = dict(list(continuous_cols.items()) + list(categorical_cols.items()))\n    \n    if training == True:\n        # Converts the label column into a constant Tensor.\n        label = tf.constant(data_set[LABEL].values)\n\n        # Returns the feature columns and the label.\n        return feature_cols, label\n    \n    return feature_cols\n\n# Model\nregressor = tf.contrib.learn.DNNRegressor(feature_columns = engineered_features, \n                                          activation_fn = tf.nn.relu, hidden_units=[200, 100, 50, 25, 12])","metadata":{"_cell_guid":"f3df3f0c-9d5a-460b-9913-bd5768c3c665","_uuid":"8069ef0a0ba6e38f4139e3b376e4e4de5546816d","execution":{"iopub.status.busy":"2022-08-01T10:56:08.562375Z","iopub.execute_input":"2022-08-01T10:56:08.562742Z","iopub.status.idle":"2022-08-01T10:56:08.668661Z","shell.execute_reply.started":"2022-08-01T10:56:08.562639Z","shell.execute_reply":"2022-08-01T10:56:08.667868Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"categorical_cols = {k: tf.SparseTensor(indices=[[i, 0] for i in range(training_set[k].size)], values = training_set[k].values, dense_shape = [training_set[k].size, 1]) for k in FEATURES_CAT}","metadata":{"_cell_guid":"b009c3cd-9048-4418-b73d-322268ddfe10","_uuid":"d7f907a9fc08d8af967c1f3f7acce50c722889fb","execution":{"iopub.status.busy":"2022-08-01T10:57:18.314153Z","iopub.execute_input":"2022-08-01T10:57:18.314538Z","iopub.status.idle":"2022-08-01T10:57:18.803247Z","shell.execute_reply.started":"2022-08-01T10:57:18.314471Z","shell.execute_reply":"2022-08-01T10:57:18.802422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Deep Neural Network Regressor with the training set which contain the data split by train test split\nregressor.fit(input_fn = lambda: input_fn_new(training_set) , steps=2000)","metadata":{"_cell_guid":"874777ec-071a-4e2f-9e60-dcbc72487487","_uuid":"71daf4fea970ed7a9232c67b5664fdaaf276203b","execution":{"iopub.status.busy":"2022-08-01T10:58:25.513244Z","iopub.execute_input":"2022-08-01T10:58:25.513574Z","iopub.status.idle":"2022-08-01T11:00:06.285505Z","shell.execute_reply.started":"2022-08-01T10:58:25.513515Z","shell.execute_reply":"2022-08-01T11:00:06.284493Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ev = regressor.evaluate(input_fn=lambda: input_fn_new(testing_set, training = True), steps=1)","metadata":{"_cell_guid":"242a2c81-5736-4322-ae86-8ff474ac1438","_uuid":"343de6e9ea87548d71fb379d5d3455d769b1fdf8","execution":{"iopub.status.busy":"2022-08-01T11:00:16.737589Z","iopub.execute_input":"2022-08-01T11:00:16.737941Z","iopub.status.idle":"2022-08-01T11:00:22.816858Z","shell.execute_reply.started":"2022-08-01T11:00:16.737877Z","shell.execute_reply":"2022-08-01T11:00:22.815879Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"loss_score4 = ev[\"loss\"]\nprint(\"Final Loss on the testing set: {0:f}\".format(loss_score4))","metadata":{"_cell_guid":"bc5b50b5-5c58-471d-a4ac-18d5b3021ebf","_uuid":"efa77284a7415bd742573492cc86a37eeff977a2","execution":{"iopub.status.busy":"2022-08-01T11:00:29.118811Z","iopub.execute_input":"2022-08-01T11:00:29.119179Z","iopub.status.idle":"2022-08-01T11:00:29.126509Z","shell.execute_reply.started":"2022-08-01T11:00:29.119120Z","shell.execute_reply":"2022-08-01T11:00:29.125279Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <center> VIII. Predictions bis","metadata":{"_cell_guid":"81cb38d1-5448-440c-af7f-92754947a1f3","_uuid":"ef71e7267dae00cbf4bc2cd88260e46bb9088ee1"}},{"cell_type":"markdown","source":"**What is Itertools Islice?**\n\nThe islice() function is part of the itertools library, and it takes an iterable object and returns a segment from it, between the elements defined by the start and end arguments given to the function: itertools.islice(iterable, start, end)","metadata":{}},{"cell_type":"code","source":"# Predictions\ny = regressor.predict(input_fn=lambda: input_fn_new(testing_set))\npredictions = list(itertools.islice(y, testing_set.shape[0]))\npredictions = pd.DataFrame(prepro_y.inverse_transform(np.array(predictions).reshape(434,1)))","metadata":{"_cell_guid":"3c1a3d70-26b9-4725-a431-09688c61b135","_uuid":"71b34cbf12b26059ea3ddab595ddae87b1e937b4","execution":{"iopub.status.busy":"2022-08-01T11:01:21.789783Z","iopub.execute_input":"2022-08-01T11:01:21.790176Z","iopub.status.idle":"2022-08-01T11:01:26.978351Z","shell.execute_reply.started":"2022-08-01T11:01:21.790109Z","shell.execute_reply":"2022-08-01T11:01:26.977141Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"matplotlib.rc('xtick', labelsize=30) \nmatplotlib.rc('ytick', labelsize=30) \n\nfig, ax = plt.subplots(figsize=(50, 40))\n\nplt.style.use('ggplot')\nplt.plot(predictions.values, reality.values, 'ro')\nplt.xlabel('Predictions', fontsize = 30)\nplt.ylabel('Reality', fontsize = 30)\nplt.title('Predictions x Reality on dataset Test', fontsize = 30)\nax.plot([reality.min(), reality.max()], [reality.min(), reality.max()], 'k--', lw=4)\nplt.show()","metadata":{"_cell_guid":"6e1fa3c8-6bc0-4f44-88bb-e102bbc15bab","_uuid":"f202b4cc9c46c2cec55f144ad457233e2d892cfa","execution":{"iopub.status.busy":"2022-08-01T11:01:29.349420Z","iopub.execute_input":"2022-08-01T11:01:29.349747Z","iopub.status.idle":"2022-08-01T11:01:30.534801Z","shell.execute_reply.started":"2022-08-01T11:01:29.349669Z","shell.execute_reply":"2022-08-01T11:01:30.533713Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_predict = regressor.predict(input_fn=lambda: input_fn_new(testing_sub, training = False))","metadata":{"_cell_guid":"784393b4-b463-4245-9256-c14dffa30b12","_uuid":"9f13d3b60f511acbb91056cf1b12d037b60146be","execution":{"iopub.status.busy":"2022-08-01T11:01:36.758429Z","iopub.execute_input":"2022-08-01T11:01:36.758949Z","iopub.status.idle":"2022-08-01T11:01:41.152612Z","shell.execute_reply.started":"2022-08-01T11:01:36.758887Z","shell.execute_reply":"2022-08-01T11:01:41.151618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"to_submit(y_predict, \"submission_cont_categ\")","metadata":{"_cell_guid":"24bc5061-452d-4e05-9d13-97cb7aab8289","_uuid":"1c9106270fd38e149054c6fcd1576c67a2db3a48","execution":{"iopub.status.busy":"2022-08-01T11:01:48.257987Z","iopub.execute_input":"2022-08-01T11:01:48.258348Z","iopub.status.idle":"2022-08-01T11:01:49.649411Z","shell.execute_reply.started":"2022-08-01T11:01:48.258272Z","shell.execute_reply":"2022-08-01T11:01:49.648535Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <center> IX. Shallow Network","metadata":{"_cell_guid":"7e016bd8-c2ce-46f2-aabd-6ab0c319de47","_uuid":"c38f966130fa61b824e432043818618a0e849fe5"}},{"cell_type":"markdown","source":"For this part we will expolore the architecture with just one Hidden Layer with several units. The question is: How many units do you need to have a good score on the leaderboard? We will try with 1000 units with the activation function Relu.","metadata":{"_cell_guid":"a669a46b-f1a9-463d-8b1b-7e45505ad178","_uuid":"a47468d133d7d191f770c3d08334864aa74d15d1"}},{"cell_type":"code","source":"# Model\nregressor = tf.contrib.learn.DNNRegressor(feature_columns = engineered_features, \n                                          activation_fn = tf.nn.relu, hidden_units=[1000])","metadata":{"_cell_guid":"62ea6c6e-da9a-43fd-b72c-a9ce5f04330c","_uuid":"827d3c0248410ceb027cb564c8cdee55965cf0e5","execution":{"iopub.status.busy":"2022-08-01T11:02:15.870253Z","iopub.execute_input":"2022-08-01T11:02:15.870597Z","iopub.status.idle":"2022-08-01T11:02:15.878762Z","shell.execute_reply.started":"2022-08-01T11:02:15.870536Z","shell.execute_reply":"2022-08-01T11:02:15.877535Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**What are hidden units in neural network?**\n\nIn neural networks, a hidden layer is located between the input and output of the algorithm, in which the function applies weights to the inputs and directs them through an activation function as the output. In short, the hidden layers perform nonlinear transformations of the inputs entered into the network.\n\nhttps://deepai.org/machine-learning-glossary-and-terms/hidden-layer-machine-learning\n","metadata":{}},{"cell_type":"code","source":"# Deep Neural Network Regressor with the training set which contain the data split by train test split\nregressor.fit(input_fn = lambda: input_fn_new(training_set) , steps=2000)","metadata":{"_cell_guid":"d06dc911-71e7-458a-9a6b-92ae6556fe77","_uuid":"a130ce41bfa507dcd3171eb97f1e56f559d9e444","execution":{"iopub.status.busy":"2022-08-01T11:04:45.177360Z","iopub.execute_input":"2022-08-01T11:04:45.177728Z","iopub.status.idle":"2022-08-01T11:06:25.264926Z","shell.execute_reply.started":"2022-08-01T11:04:45.177652Z","shell.execute_reply":"2022-08-01T11:06:25.263774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ev = regressor.evaluate(input_fn=lambda: input_fn_new(testing_set, training = True), steps=1)\nloss_score5 = ev[\"loss\"]","metadata":{"_cell_guid":"a071d27f-5257-4f3c-bccf-5e9a80daed0f","_uuid":"352fd136f75cab9e514d3fa8918385bb3ce768fb","execution":{"iopub.status.busy":"2022-08-01T11:10:41.336834Z","iopub.execute_input":"2022-08-01T11:10:41.337229Z","iopub.status.idle":"2022-08-01T11:10:46.451433Z","shell.execute_reply.started":"2022-08-01T11:10:41.337166Z","shell.execute_reply":"2022-08-01T11:10:46.450539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Final Loss on the testing set: {0:f}\".format(loss_score5))","metadata":{"_cell_guid":"3c76cf13-b147-414c-a2d0-4eee1b62b297","_uuid":"e5063613e9925962428ed861437449f458ff9d29","execution":{"iopub.status.busy":"2022-08-01T11:10:49.183615Z","iopub.execute_input":"2022-08-01T11:10:49.184035Z","iopub.status.idle":"2022-08-01T11:10:49.189924Z","shell.execute_reply.started":"2022-08-01T11:10:49.183965Z","shell.execute_reply":"2022-08-01T11:10:49.188728Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_predict = regressor.predict(input_fn=lambda: input_fn_new(testing_sub, training = False))    \nto_submit(y_predict, \"submission_shallow\")","metadata":{"_cell_guid":"58bc0c60-456c-408b-a6d7-255a4506db36","_uuid":"87be0f6006d75564020d7967528cec97a6664e85","execution":{"iopub.status.busy":"2022-08-01T11:10:51.824741Z","iopub.execute_input":"2022-08-01T11:10:51.825089Z","iopub.status.idle":"2022-08-01T11:10:59.014017Z","shell.execute_reply.started":"2022-08-01T11:10:51.825029Z","shell.execute_reply":"2022-08-01T11:10:59.012854Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <center> X. Conclusion","metadata":{"_cell_guid":"51ab5586-53f3-43e0-926b-e178f856f5a1","_uuid":"ee22d288772717d0f61fa1ca8db79b8f765ecccc"}},{"cell_type":"code","source":"list_score = [loss_score1, loss_score2, loss_score3, loss_score4,loss_score5]\nlist_model = ['Relu_cont', 'LRelu_cont', 'Elu_cont', 'Relu_cont_categ','Shallow_1ku']","metadata":{"_cell_guid":"9547c21a-4f23-49cf-8cc6-269c328b8ec5","_uuid":"548cadf701cde3e7850171dd6f52a24fd42add0f","execution":{"iopub.status.busy":"2022-08-01T11:10:59.017178Z","iopub.execute_input":"2022-08-01T11:10:59.017510Z","iopub.status.idle":"2022-08-01T11:10:59.025760Z","shell.execute_reply.started":"2022-08-01T11:10:59.017449Z","shell.execute_reply":"2022-08-01T11:10:59.024909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list_score","metadata":{"execution":{"iopub.status.busy":"2022-08-01T11:11:00.895827Z","iopub.execute_input":"2022-08-01T11:11:00.896142Z","iopub.status.idle":"2022-08-01T11:11:00.903114Z","shell.execute_reply.started":"2022-08-01T11:11:00.896083Z","shell.execute_reply":"2022-08-01T11:11:00.901879Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt; plt.rcdefaults()\n\nplt.style.use('ggplot')\nobjects = list_model #['Relu_cont', 'LRelu_cont', 'Elu_cont', 'Relu_cont_categ','Shallow_1ku']\ny_pos = np.arange(len(objects)) #array([0, 1, 2, 3, 4])\nperformance = list_score #[0.0016870112, 0.001616128, 0.0020051321, 0.002015396, 0.0018025363]\n \nplt.barh(y_pos, performance, align='center', alpha=0.9)\nplt.yticks(y_pos, objects)\nplt.xlabel('Loss ')\nplt.title('Model compared without hypertuning')\n \nplt.show()","metadata":{"_cell_guid":"5b1f7fb1-f8ba-4176-9bc6-a7c85d520e50","_uuid":"38ac30fe3da8cd6a6bd98e78c8aaadbfa16ccb7a","execution":{"iopub.status.busy":"2022-08-01T11:11:54.905017Z","iopub.execute_input":"2022-08-01T11:11:54.905373Z","iopub.status.idle":"2022-08-01T11:11:55.101592Z","shell.execute_reply.started":"2022-08-01T11:11:54.905293Z","shell.execute_reply":"2022-08-01T11:11:55.100675Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pos","metadata":{"execution":{"iopub.status.busy":"2022-08-01T11:12:04.176471Z","iopub.execute_input":"2022-08-01T11:12:04.176842Z","iopub.status.idle":"2022-08-01T11:12:04.183169Z","shell.execute_reply.started":"2022-08-01T11:12:04.176778Z","shell.execute_reply":"2022-08-01T11:12:04.181869Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"regressor.get_params","metadata":{"execution":{"iopub.status.busy":"2022-08-01T11:12:07.775553Z","iopub.execute_input":"2022-08-01T11:12:07.776042Z","iopub.status.idle":"2022-08-01T11:12:07.783629Z","shell.execute_reply.started":"2022-08-01T11:12:07.775966Z","shell.execute_reply":"2022-08-01T11:12:07.782419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"regressor.evaluate","metadata":{"execution":{"iopub.status.busy":"2022-08-01T11:12:19.296367Z","iopub.execute_input":"2022-08-01T11:12:19.296930Z","iopub.status.idle":"2022-08-01T11:12:19.304990Z","shell.execute_reply.started":"2022-08-01T11:12:19.296722Z","shell.execute_reply":"2022-08-01T11:12:19.303804Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"regressor.model_fn","metadata":{"execution":{"iopub.status.busy":"2022-08-01T11:12:27.904613Z","iopub.execute_input":"2022-08-01T11:12:27.904964Z","iopub.status.idle":"2022-08-01T11:12:27.911952Z","shell.execute_reply.started":"2022-08-01T11:12:27.904905Z","shell.execute_reply":"2022-08-01T11:12:27.910890Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So, I hope that this small introduction will be useful ! With this code you can **build a regression model with Tensorflow with continuous and categorical features plus add a new activation function (LeakyRelu).**\n\nIf you want to improve the results you can re-build the models on the whole of data. You can see that I'm used just 67% of the training set to build my models.\n\nTake my code and play with it: More Hyperparameters, 100% of the training set to build the next models, try with other activation function etc...\n","metadata":{"_cell_guid":"9b48d35b-9281-48de-8534-b0448b99918b","_uuid":"285871604d03aa1b18672e871c4641b0098a11a1"}}]}