{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# __House Prices - Advanced Regression Techniques__","metadata":{}},{"cell_type":"markdown","source":"I am not a expert in data science. There might be some issues. I hope you like this. These all written by myself and please leave a comment.. ","metadata":{}},{"cell_type":"markdown","source":"<img src='https://storage.googleapis.com/kaggle-competitions/kaggle/5407/media/housesbanner.png'>","metadata":{}},{"cell_type":"markdown","source":"## __Workflow Stages__","metadata":{}},{"cell_type":"markdown","source":"1. Problem definition.\n2. Acquire training and testing data.\n3. Wrangle, Prepare, Cleanse the data.\n4. Analyze, identify patterns and explore the data.\n5. Model, Predict and solve the problem.\n6. Visualize, report, and present the problem solving steps and final solution.\n7. Supply or submit the results.","metadata":{}},{"cell_type":"markdown","source":"The workflow indicates general sequence of how each stage may follow the other. However there are use cases with exceptions.","metadata":{}},{"cell_type":"markdown","source":"* We may combine multiple workflow stages. We may analyze by visualizing data.\n* Perform a stage earlier than indicated. We may analyze data before and after wrangling.\n* Perform a stage multiple times in our workflow. Visualize stage may be used multiple times.\n* We may drop a stage altogether","metadata":{}},{"cell_type":"markdown","source":"## __Problem definition__","metadata":{}},{"cell_type":"markdown","source":"__Predict sales prices and practice feature engineering, RFs, and gradient boosting__","metadata":{}},{"cell_type":"markdown","source":"Ask a home buyer to describe their dream house, and they probably won't begin with the height of the basement ceiling or the proximity to an east-west railroad. But this playground competition's dataset proves that much more influences price negotiations than the number of bedrooms or a white-picket fence.\n\nWith 79 explanatory variables describing (almost) every aspect of residential homes in Ames, Iowa, this competition challenges you to predict the final price of each home.","metadata":{}},{"cell_type":"markdown","source":"### __Workflow goals__","metadata":{}},{"cell_type":"markdown","source":"The datascience solutions workflow solves for seven major goals.","metadata":{}},{"cell_type":"markdown","source":"__Classifing__. We may want to classify or categorize our samples. We may also want to undestand the implications or correlations of different classes with ourr solution goals\n\n__Correlation__. One can approach the problem based on available features within the training dataset. Which features within the dataset contribute significantly to our solution goal? Statistically speaking is there a correlation among a feature and solution goal? As the feature values change does the solution state change as well, and visa-versa? This can be tested both for numerical and categorical features in the given dataset. We may also want to determine correlation among features other than survival for subsequent goals and workflow stages. Correlating certain features may help in creating, completing, or correcting features.\n\n__Converting__. For modelling stage, one needs to prepare the data. Depending on the choice of model algorithm one may require all feature to be converted to numerical equivalent values. So for instance converting text categorical values to numerical values.\n\n__Completing__. Data preparation may also require us to estimate any missing values within a feature. Model algorithm may work best when there are no missing values.\n\n__Correcting__. We may also analyze the given training dataset for errors or possibly innacurate values within features and try to corrent these values or exclude the samples containing the errors. One way to do this is to detect any outliers among our samples or features. We may also completely discard a feature if it is not contributing to the analysis or may significantly skew the result.\n\n__Creating__. Can we create new feature based on an existing feature or a set of features, such that the new feature follows the correlation, conversion, completeness goals.\n\n__Charting__. How to slect the right visualization plots and charts depending on nature of the data and the solution goals.","metadata":{}},{"cell_type":"code","source":"# data analysis and wrangling\nimport pandas as pd\nimport numpy as np\nimport random as rnd\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.preprocessing import MinMaxScaler\n\n# visualization\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n%matplotlib inline\nfrom tabulate import tabulate\nimport itertools\n\n# model training\nfrom sklearn.ensemble import RandomForestRegressor, GradientBoostingRegressor, AdaBoostRegressor, BaggingRegressor\nfrom sklearn.kernel_ridge import KernelRidge\nfrom sklearn.linear_model import Ridge, RidgeCV\nfrom sklearn.linear_model import ElasticNet, ElasticNetCV\nfrom sklearn.svm import SVR\nfrom mlxtend.regressor import StackingCVRegressor\nimport lightgbm as lgb\nfrom lightgbm import LGBMRegressor\nfrom xgboost import XGBRegressor\nfrom sklearn.model_selection import GridSearchCV\nfrom sklearn.tree import DecisionTreeRegressor","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:52:30.573965Z","iopub.execute_input":"2022-08-06T06:52:30.574383Z","iopub.status.idle":"2022-08-06T06:52:30.587822Z","shell.execute_reply.started":"2022-08-06T06:52:30.574347Z","shell.execute_reply":"2022-08-06T06:52:30.586424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## __Acquire the data__","metadata":{}},{"cell_type":"markdown","source":"Import train and test data sets using python pandas library and we create a variable called combine with combining train and test data sets together.","metadata":{}},{"cell_type":"code","source":"# Training data\ntrain_df = pd.read_csv('../input/house-prices-advanced-regression-techniques/train.csv')\n# Test data\ntest_df = pd.read_csv('../input/house-prices-advanced-regression-techniques/test.csv')\n# We combine the training and testing data for our convenience\ncombine = [train_df, test_df]","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:52:30.590065Z","iopub.execute_input":"2022-08-06T06:52:30.591081Z","iopub.status.idle":"2022-08-06T06:52:30.638824Z","shell.execute_reply.started":"2022-08-06T06:52:30.591026Z","shell.execute_reply":"2022-08-06T06:52:30.637629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## __Analyze data by describing__","metadata":{}},{"cell_type":"markdown","source":"With pandas library we can use it to answer following questions.","metadata":{}},{"cell_type":"markdown","source":"#### __Which features are available in the dataset?__","metadata":{}},{"cell_type":"markdown","source":"In training dataset there are 81 features and testing dataset there are 80 features.","metadata":{}},{"cell_type":"code","source":"train_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:52:30.640220Z","iopub.execute_input":"2022-08-06T06:52:30.640662Z","iopub.status.idle":"2022-08-06T06:52:30.647857Z","shell.execute_reply.started":"2022-08-06T06:52:30.640628Z","shell.execute_reply":"2022-08-06T06:52:30.646656Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:52:30.650908Z","iopub.execute_input":"2022-08-06T06:52:30.652199Z","iopub.status.idle":"2022-08-06T06:52:30.661192Z","shell.execute_reply.started":"2022-08-06T06:52:30.652148Z","shell.execute_reply":"2022-08-06T06:52:30.660091Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train_df.columns.values)","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:52:30.662781Z","iopub.execute_input":"2022-08-06T06:52:30.663442Z","iopub.status.idle":"2022-08-06T06:52:30.671287Z","shell.execute_reply.started":"2022-08-06T06:52:30.663405Z","shell.execute_reply":"2022-08-06T06:52:30.670245Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This is a long dataset. To avoid some columns hiding set the max column display to none","metadata":{}},{"cell_type":"code","source":"pd.set_option('display.max_columns', None)\npd.set_option('display.max_rows', None)","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:52:30.673391Z","iopub.execute_input":"2022-08-06T06:52:30.673998Z","iopub.status.idle":"2022-08-06T06:52:30.681755Z","shell.execute_reply.started":"2022-08-06T06:52:30.673957Z","shell.execute_reply":"2022-08-06T06:52:30.680766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### __Which features are categorical?__","metadata":{}},{"cell_type":"markdown","source":"These values classify the samples into sets of similar samples. (nominal, ordinal, ratio, interval based)","metadata":{}},{"cell_type":"markdown","source":"* Categorical\n    * MSZoning\n    * Street\n    * Alley\n    * LotShape\n    * LandContour\n    * Utilities\n    * LongConfig\n    * LandSlope\n    * Neighbhorhood\n    * Condition1\n    * Condition2\n    * BldgType\n    * HouseStyle\n    * RoofStyle\n    * RoofMalt\n    * Exterior1st\n    * Exterior2nd\n    * MasVnrType\n    * ExterQual\n    * ExterCond\n    * Foundation\n    * BsmtQual\n    * BsmtCond\n    * BsmtExposure\n    * BsmtFinType1\n    * BsmtFinType2\n    * Heating\n    * HeatingQC\n    * CentralAir\n    * Electrical\n    * KichenQual\n    * Functional\n    * FireplaceQu\n    * GarageType\n    * GarageFinish\n    * GarageQual\n    * GarageCond\n    * PavedDrive\n    * PoolQC\n    * Fence\n    * MiscFeature\n    * SaleType\n    * SaleCondition\n* Ordinal\n    * YearBuilt\n    * YearRemodAdd\n    * BsmtFullBath\n    * BsmtHalfBath\n    * FullBath\n    * HalfBath\n    * BedroomAbvGr\n    * KitchenAbvGr\n    * Fireplaces\n    * GarageYrBlt\n    * GarageCars\n    \n\n","metadata":{}},{"cell_type":"markdown","source":"#### __Which features are numerical?__","metadata":{}},{"cell_type":"markdown","source":"These values change from sample to sample. (Discrete, Continuous, Timeseries based)","metadata":{}},{"cell_type":"markdown","source":"* Continuous\n    * Id\n    * MSSubClass\n    * LotFrontage\n    * LotArea\n    * MasVnrArea\n    * BsmtFinSF1\n    * BsmtFinSF2\n    * BsmtUnfSF\n    * TotalBsmtSF\n    * 1stFlrSF\n    * 2ndFlrSF\n    * GrLivArea \n    * GarageArea\n    * WoodDeckSF\n    * OpenPorchSF\n    * EnclosedPorch\n    * 3SsnPorch\n    * PoolArea\n    * MiscVal\n    * MoSold\n    * SalePrice\n\n* Discrete\n    * OverallQual\n    * OverallCond\n    * LowQualFinSF\n    * ScreenPorch\n\n","metadata":{}},{"cell_type":"markdown","source":"#### __Which features are mixed type?__","metadata":{}},{"cell_type":"markdown","source":"Numerical and alphanumerical data within the same feature","metadata":{}},{"cell_type":"markdown","source":"* Alphanumerical\n    * LotShape\n    * LotConfig\n    * BldgType","metadata":{}},{"cell_type":"markdown","source":"#### __Which feature contain blank, null, empty values?__","metadata":{}},{"cell_type":"markdown","source":"* In train dataset:\n    * LotFrontage       259\n    * Alley            1369\n    * BsmtQual           37\n    * BsmtCond           37\n    * BsmtExposure       38\n    * BsmtFinType1       37\n    * BsmtFinType2       38\n    * Electrical          1\n    * FireplaceQu       690\n    * GarageType         81\n    * GarageYrBlt        81\n    * GarageFinish       81\n    * GarageQual         81\n    * GarageCond         81\n    * PoolQC           1453\n    * Fence            1179\n    * MiscFeature      1406\n* In test dataset\n    * MSZoning            4\n    * LotFrontage       227\n    * Alley            1352\n    * Utilities           2\n    * Exterior1st         1\n    * Exterior2nd         1\n    * MasVnrType         16\n    * MasVnrArea         15\n    * BsmtQual           44\n    * BsmtCond           45\n    * BsmtExposure       44\n    * BsmtFinType1       42\n    * BsmtFinSF1          1\n    * BsmtFinType2       42\n    * BsmtFinSF2          1\n    * BsmtUnfSF           1\n    * TotalBsmtSF         1\n    * BsmtFullBath        2\n    * BsmtHalfBath        2\n    * KitchenQual         1\n    * Functional          2\n    * FireplaceQu       730\n    * GarageType         76\n    * GarageYrBlt        78\n    * GarageFinish       78\n    * GarageCars          1\n    * GarageArea          1\n    * GarageQual         78\n    * GarageCond         78\n    * PoolQC           1456\n    * Fence            1169\n    * MiscFeature      1408\n    * SaleType            1","metadata":{}},{"cell_type":"code","source":"train_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:52:30.683186Z","iopub.execute_input":"2022-08-06T06:52:30.683795Z","iopub.status.idle":"2022-08-06T06:52:30.702313Z","shell.execute_reply.started":"2022-08-06T06:52:30.683761Z","shell.execute_reply":"2022-08-06T06:52:30.701187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:52:30.705255Z","iopub.execute_input":"2022-08-06T06:52:30.706216Z","iopub.status.idle":"2022-08-06T06:52:30.718913Z","shell.execute_reply.started":"2022-08-06T06:52:30.706180Z","shell.execute_reply":"2022-08-06T06:52:30.717546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### __What are the data types in various features?__","metadata":{}},{"cell_type":"markdown","source":"* In training dataset there are 43 object(string), 35 int, 3 float64 data types.\n* In testing dataset there are 43 object(sting), 26 int, 11 float64 data types.","metadata":{}},{"cell_type":"code","source":"train_df.dtypes.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:52:30.722203Z","iopub.execute_input":"2022-08-06T06:52:30.723080Z","iopub.status.idle":"2022-08-06T06:52:30.733481Z","shell.execute_reply.started":"2022-08-06T06:52:30.723004Z","shell.execute_reply":"2022-08-06T06:52:30.732298Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.dtypes.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:52:30.734881Z","iopub.execute_input":"2022-08-06T06:52:30.735578Z","iopub.status.idle":"2022-08-06T06:52:30.745704Z","shell.execute_reply.started":"2022-08-06T06:52:30.735496Z","shell.execute_reply":"2022-08-06T06:52:30.744439Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### __What is the distribution of numerical feature values in the dataset?__","metadata":{}},{"cell_type":"markdown","source":"* This dataset has 1460 samples. \n* The max quality is 10 and min quality is 1\n* The houses in this dataset belongs to 1870-2010 build years.\n* Max bathrooms count is 3 in these houses.\n* The oldest garage was build in 1900. \n* Most houses don't have a pool\n* First house was sold in 2006\n","metadata":{}},{"cell_type":"code","source":"train_df.describe()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:52:30.747477Z","iopub.execute_input":"2022-08-06T06:52:30.747867Z","iopub.status.idle":"2022-08-06T06:52:30.874004Z","shell.execute_reply.started":"2022-08-06T06:52:30.747823Z","shell.execute_reply":"2022-08-06T06:52:30.872816Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### __What is the distribution of categorical features?__","metadata":{}},{"cell_type":"markdown","source":"* Most houses in this dataset access the road in Pave type\n* Most houses are normal condition\n* About half of houses in this dataset belongs to 1Story style\n* A few amount of houses have a fence\n* Sale condition is normal in most of houses\n","metadata":{}},{"cell_type":"code","source":"train_df.describe(include=['O'])","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:52:30.875568Z","iopub.execute_input":"2022-08-06T06:52:30.875912Z","iopub.status.idle":"2022-08-06T06:52:30.981096Z","shell.execute_reply.started":"2022-08-06T06:52:30.875880Z","shell.execute_reply":"2022-08-06T06:52:30.979588Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### __Assumptions based on data analysis__","metadata":{}},{"cell_type":"markdown","source":"We arrived following assumtions based on data analysis so far.","metadata":{}},{"cell_type":"markdown","source":"* __Correlation__: We want to know how well does each feature correlate with SalePrice.","metadata":{}},{"cell_type":"markdown","source":"* __Completing:__ We want to complete all the empty values in the dataset\n","metadata":{}},{"cell_type":"markdown","source":"## __Analyze by pivolating features?__","metadata":{}},{"cell_type":"markdown","source":"We can analyze fureture correlation by pivolating each feature.","metadata":{}},{"cell_type":"markdown","source":"* We can see the variation of SalePrice with different features with different categories. So we should use the all these features for our model training.","metadata":{}},{"cell_type":"code","source":"# first create a list with categorical features\ncat_futures = ['MSZoning',\t'Street',\t'Alley',\t'LotShape',\t'LandContour',\t'Utilities'\t,'LotConfig'\t,'LandSlope',\t'Neighborhood'\t,'Condition1'\t,'Condition2',\t'BldgType',\t'HouseStyle'\t,'RoofStyle',\t'RoofMatl'\t,'Exterior1st',\t'Exterior2nd',\t'MasVnrType'\t,'ExterQual',\t'ExterCond'\t,'Foundation',\t'BsmtQual'\t,'BsmtCond',\t'BsmtExposure',\t'BsmtFinType1'\t,'BsmtFinType2',\t'Heating',\t'HeatingQC',\t'CentralAir',\t'Electrical'\t,'KitchenQual',\t'Functional'\t,'FireplaceQu','GarageType',\t'GarageFinish',\t'GarageQual'\t,'GarageCond',\t'PavedDrive'\t,'PoolQC',\t'Fence',\t'MiscFeature',\t'SaleType',\t'SaleCondition']","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:52:30.983063Z","iopub.execute_input":"2022-08-06T06:52:30.983712Z","iopub.status.idle":"2022-08-06T06:52:30.991942Z","shell.execute_reply.started":"2022-08-06T06:52:30.983675Z","shell.execute_reply":"2022-08-06T06:52:30.990367Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for each in cat_futures:\n    #If you don't have tabulate try commented code line\n    # print(train_df[[each,'SalePrice']].groupby([each], as_index=False).mean().sort_values(by='SalePrice', ascending=False),\"\\n ------------------------\")\n    print(tabulate((train_df[[each,'SalePrice']].groupby([each], as_index=False).mean().sort_values(by='SalePrice', ascending=False)),headers='keys', tablefmt='fancy_grid', showindex=False))","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:52:30.993879Z","iopub.execute_input":"2022-08-06T06:52:30.994363Z","iopub.status.idle":"2022-08-06T06:52:31.168191Z","shell.execute_reply.started":"2022-08-06T06:52:30.994316Z","shell.execute_reply":"2022-08-06T06:52:31.166907Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## __Analize by visualizing data__","metadata":{}},{"cell_type":"markdown","source":"We can continue our assumptions using visualizing data.","metadata":{}},{"cell_type":"markdown","source":"#### __Correlation numerical features__","metadata":{}},{"cell_type":"markdown","source":"* __Observations__\n    * We can see the SalePrice is incresing with getting higher each feture values\n    * Also SalePrice is getting higher with some categorical numeric feature values ","metadata":{}},{"cell_type":"markdown","source":"* __Decisions__\n    * Use each feature for model training\n    * There are some high range values. So we want to add scaler for each values","metadata":{}},{"cell_type":"markdown","source":"First We want to create two categorical and numerical variable","metadata":{}},{"cell_type":"code","source":"cat = train_df.select_dtypes('object')\nnum = train_df.select_dtypes('number')\n# But we dont want id and SellPrice columns. Lets drop them\nnum = num.drop(['Id', 'SalePrice'], axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:52:31.169814Z","iopub.execute_input":"2022-08-06T06:52:31.170180Z","iopub.status.idle":"2022-08-06T06:52:31.181190Z","shell.execute_reply.started":"2022-08-06T06:52:31.170148Z","shell.execute_reply":"2022-08-06T06:52:31.179819Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"color_cycle= itertools.cycle([\"orange\",\"pink\",\"blue\",\"brown\",\"red\",\"grey\",\"yellow\",\"green\"])\nfor i in num:\n    plt.scatter(train_df[i],train_df['SalePrice'],color=next(color_cycle))\n    plt.title(i)\n    plt.xlabel(i)\n    plt.ylabel('SalePrice')\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:52:31.182892Z","iopub.execute_input":"2022-08-06T06:52:31.183977Z","iopub.status.idle":"2022-08-06T06:52:39.463415Z","shell.execute_reply.started":"2022-08-06T06:52:31.183938Z","shell.execute_reply":"2022-08-06T06:52:39.462427Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### __Correlation Categorical feature values__","metadata":{}},{"cell_type":"markdown","source":"* __Observations__:\n    * We can see how sale price different in each category\n    * Some feature categories have few samples\n    \n* __Desicions__:\n    * Use these categorical feature for model training","metadata":{}},{"cell_type":"code","source":"for i in cat:\n    plt.figure(figsize=(20,4))\n    plt.subplots_adjust(hspace=.25)\n    plt.subplot(1,2,1)\n    plt.title('How price is different')\n    sns.barplot(x=i, y='SalePrice',data=train_df)\n    plt.subplot(1,2,2)\n    plt.title('Home much value counts')\n    sns.stripplot(x=i, y='SalePrice',data=train_df)","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:52:39.464627Z","iopub.execute_input":"2022-08-06T06:52:39.465161Z","iopub.status.idle":"2022-08-06T06:53:00.458670Z","shell.execute_reply.started":"2022-08-06T06:52:39.465128Z","shell.execute_reply":"2022-08-06T06:53:00.457274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## __Wrangle data__","metadata":{}},{"cell_type":"markdown","source":"We have collected several assumptions and decisions regarding our datasets and solution requrements. So far we did not change any value or any feature. Lets execute our decisions and assumptions for correcting, creating, and completing goals.","metadata":{}},{"cell_type":"markdown","source":"### __Completing empty values__","metadata":{}},{"cell_type":"markdown","source":"We should complete empty values in each feature. ","metadata":{}},{"cell_type":"code","source":"for dataset in combine:\n    dataset['LotFrontage'] = dataset['LotFrontage'].fillna(train_df['LotFrontage'].median())\n    dataset['LotFrontage'] = dataset['LotFrontage'].fillna(train_df['LotFrontage'].median())\n    dataset['Alley'] = dataset['Alley'].fillna('None')\n    dataset['BsmtQual'] = dataset['BsmtQual'].fillna('NoBsmt')\n    dataset['BsmtCond'] = dataset['BsmtCond'].fillna('NoBsmt')\n    dataset['BsmtExposure'] = dataset['BsmtExposure'].fillna('NoBsmt')\n    dataset['BsmtFinType1'] = dataset['BsmtFinType1'].fillna('NoBsmt')\n    dataset['BsmtFinType2'] = dataset['BsmtFinType2'].fillna('NoBsmt')\n    dataset['Electrical'] = dataset['Electrical'].fillna(train_df.Electrical.dropna().mode()[0])\n    dataset['FireplaceQu'] = dataset['FireplaceQu'].fillna('Gd')\n    dataset['GarageType'] = dataset['GarageType'].fillna('NoGarage')\n    dataset['GarageYrBlt'] = dataset['GarageYrBlt'].fillna(0)\n    dataset['GarageFinish'] = dataset['GarageFinish'].fillna('NoGarage')\n    dataset['GarageQual'] = dataset['GarageQual'].fillna('NoGarage')\n    dataset['GarageCond'] = dataset['GarageCond'].fillna('NoGarage')\n    dataset['PoolQC'] = dataset['PoolQC'].fillna('Normal')\n    dataset['Fence'] = dataset['Fence'].fillna('NoFence')\n    dataset['MiscFeature'] = dataset['MiscFeature'].fillna('None')\n    dataset['MSZoning'] = dataset['MSZoning'].fillna('RL')\n    dataset['Utilities'] = dataset['Utilities'].fillna('NoSeWa')\n    dataset['Exterior1st'] = dataset['Exterior1st'].fillna(train_df.Exterior1st.dropna().mode()[0])\n    dataset['Exterior2nd'] = dataset['Exterior2nd'].fillna('Other')\n    dataset['MasVnrType'] = dataset['MasVnrType'].fillna('None')\n    dataset['MasVnrArea'] = dataset['MasVnrArea'].fillna(0)\n    dataset['BsmtFinSF1'] = dataset['BsmtFinSF1'].fillna(0)\n    dataset['BsmtFinSF2'] = dataset['BsmtFinSF2'].fillna(0)\n    dataset['BsmtUnfSF'] = dataset['BsmtUnfSF'].fillna(0)\n    dataset['TotalBsmtSF'] = dataset['TotalBsmtSF'].fillna(0)\n    dataset['BsmtFullBath'] = dataset['BsmtFullBath'].fillna(0)\n    dataset['BsmtHalfBath'] = dataset['BsmtHalfBath'].fillna(0)\n    dataset['KitchenQual'] = dataset['KitchenQual'].fillna('Gd')\n    dataset['Functional'] = dataset['Functional'].fillna('Typ')\n    dataset['GarageCars'] = dataset['GarageCars'].fillna(0)\n    dataset['GarageArea'] = dataset['GarageArea'].fillna(0)\n    dataset['SaleType'] = dataset['SaleType'].fillna('Oth')\n\ncombine = [train_df,test_df]","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:53:00.465975Z","iopub.execute_input":"2022-08-06T06:53:00.466725Z","iopub.status.idle":"2022-08-06T06:53:00.518350Z","shell.execute_reply.started":"2022-08-06T06:53:00.466688Z","shell.execute_reply":"2022-08-06T06:53:00.517019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.isnull().sum().value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:53:00.519594Z","iopub.execute_input":"2022-08-06T06:53:00.520024Z","iopub.status.idle":"2022-08-06T06:53:00.533422Z","shell.execute_reply.started":"2022-08-06T06:53:00.519989Z","shell.execute_reply":"2022-08-06T06:53:00.532348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.isnull().sum().value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:53:00.535092Z","iopub.execute_input":"2022-08-06T06:53:00.535930Z","iopub.status.idle":"2022-08-06T06:53:00.549731Z","shell.execute_reply.started":"2022-08-06T06:53:00.535894Z","shell.execute_reply":"2022-08-06T06:53:00.548758Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we can see there isn't any null values.","metadata":{}},{"cell_type":"markdown","source":"### __Converting categorical features to numerical__","metadata":{}},{"cell_type":"code","source":"encoder = LabelEncoder()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:53:00.551143Z","iopub.execute_input":"2022-08-06T06:53:00.551769Z","iopub.status.idle":"2022-08-06T06:53:00.556418Z","shell.execute_reply.started":"2022-08-06T06:53:00.551734Z","shell.execute_reply":"2022-08-06T06:53:00.555362Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in cat:\n    for dataset in combine:\n        dataset[i] = encoder.fit_transform(dataset[i])\ncombine = [train_df,test_df]","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:53:00.558560Z","iopub.execute_input":"2022-08-06T06:53:00.559934Z","iopub.status.idle":"2022-08-06T06:53:00.645317Z","shell.execute_reply.started":"2022-08-06T06:53:00.559844Z","shell.execute_reply":"2022-08-06T06:53:00.644348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:53:00.647263Z","iopub.execute_input":"2022-08-06T06:53:00.647742Z","iopub.status.idle":"2022-08-06T06:53:00.693476Z","shell.execute_reply.started":"2022-08-06T06:53:00.647695Z","shell.execute_reply":"2022-08-06T06:53:00.692273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:53:00.694758Z","iopub.execute_input":"2022-08-06T06:53:00.695085Z","iopub.status.idle":"2022-08-06T06:53:00.743370Z","shell.execute_reply.started":"2022-08-06T06:53:00.695055Z","shell.execute_reply":"2022-08-06T06:53:00.742166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### __Drop unwanted features__","metadata":{}},{"cell_type":"markdown","source":"Now we want to drop id and SalePrice columns. But before dropping lets create a copy of datasets. ","metadata":{}},{"cell_type":"code","source":"train_df_copy = train_df.copy()\ntest_df_copy = test_df.copy()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:53:00.745010Z","iopub.execute_input":"2022-08-06T06:53:00.745499Z","iopub.status.idle":"2022-08-06T06:53:00.754003Z","shell.execute_reply.started":"2022-08-06T06:53:00.745439Z","shell.execute_reply":"2022-08-06T06:53:00.752987Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = train_df.drop(['Id', 'SalePrice'], axis=1)\ntest_df = test_df.drop('Id', axis=1)\ncombine = [train_df,test_df]","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:53:00.755122Z","iopub.execute_input":"2022-08-06T06:53:00.756019Z","iopub.status.idle":"2022-08-06T06:53:00.774831Z","shell.execute_reply.started":"2022-08-06T06:53:00.755982Z","shell.execute_reply":"2022-08-06T06:53:00.773407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## __Scaling__","metadata":{}},{"cell_type":"markdown","source":"There are huge value ranges in the dataset. If we use this data to train our model might be a issue. So we want to scale these values before training.","metadata":{}},{"cell_type":"markdown","source":"I'm using here min max scaler.","metadata":{}},{"cell_type":"code","source":"scaler = MinMaxScaler()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:53:00.776541Z","iopub.execute_input":"2022-08-06T06:53:00.777044Z","iopub.status.idle":"2022-08-06T06:53:00.784219Z","shell.execute_reply.started":"2022-08-06T06:53:00.776999Z","shell.execute_reply":"2022-08-06T06:53:00.782664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"columns = train_df.columns.values","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:53:00.786370Z","iopub.execute_input":"2022-08-06T06:53:00.786929Z","iopub.status.idle":"2022-08-06T06:53:00.796638Z","shell.execute_reply.started":"2022-08-06T06:53:00.786892Z","shell.execute_reply":"2022-08-06T06:53:00.795584Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for dataset in combine:\n        dataset_scaled = scaler.fit_transform(dataset[columns])\n        dataset[columns] = dataset_scaled\ncombine = [train_df,test_df]","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:53:00.798886Z","iopub.execute_input":"2022-08-06T06:53:00.799269Z","iopub.status.idle":"2022-08-06T06:53:00.855683Z","shell.execute_reply.started":"2022-08-06T06:53:00.799233Z","shell.execute_reply":"2022-08-06T06:53:00.854516Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:53:00.857142Z","iopub.execute_input":"2022-08-06T06:53:00.857537Z","iopub.status.idle":"2022-08-06T06:53:00.938043Z","shell.execute_reply.started":"2022-08-06T06:53:00.857500Z","shell.execute_reply":"2022-08-06T06:53:00.936783Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## __Model train, predict and slove__","metadata":{}},{"cell_type":"markdown","source":"Now our dataset is looking good and we have to train the model. Then we can use the model to slove the problem solution. There are many model algorithms to use. But our problem is classification and reggression problem in supervised learning. So we can use these models.","metadata":{}},{"cell_type":"markdown","source":"* RandomForestRegressor\n* GradientBoostingRegressor\n* AdaBoostRegressor\n* BaggingRegressor\n* KernelRidge\n* Ridge\n* RidgeCV\n* ElasticNet\n* ElasticNetCV\n* SVR\n* StackingCVRegressor\n* lgb\n* LGBMRegressor\n* XGBRegressor","metadata":{}},{"cell_type":"code","source":"X_train = train_df\ny_train = train_df_copy['SalePrice']\nX_test = test_df\nX_train.shape, y_train.shape, X_test.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:53:00.939679Z","iopub.execute_input":"2022-08-06T06:53:00.940104Z","iopub.status.idle":"2022-08-06T06:53:00.948447Z","shell.execute_reply.started":"2022-08-06T06:53:00.940069Z","shell.execute_reply":"2022-08-06T06:53:00.947524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we can train our model","metadata":{}},{"cell_type":"code","source":"# Random forest reggresor\nrndfreg = RandomForestRegressor()\nrndfreg.fit(X_train, y_train)\nprint(round(rndfreg.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:53:00.950396Z","iopub.execute_input":"2022-08-06T06:53:00.950820Z","iopub.status.idle":"2022-08-06T06:53:03.199882Z","shell.execute_reply.started":"2022-08-06T06:53:00.950786Z","shell.execute_reply":"2022-08-06T06:53:03.198617Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Gradient boosting\ngbreg = GradientBoostingRegressor()\ngbreg.fit(X_train, y_train)\nprint(round(gbreg.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:53:03.201385Z","iopub.execute_input":"2022-08-06T06:53:03.201987Z","iopub.status.idle":"2022-08-06T06:53:03.966518Z","shell.execute_reply.started":"2022-08-06T06:53:03.201941Z","shell.execute_reply":"2022-08-06T06:53:03.965359Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Ada boosting\nabreg = AdaBoostRegressor()\nabreg.fit(X_train, y_train)\nprint(round(abreg.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:53:03.967772Z","iopub.execute_input":"2022-08-06T06:53:03.968128Z","iopub.status.idle":"2022-08-06T06:53:04.367210Z","shell.execute_reply.started":"2022-08-06T06:53:03.968097Z","shell.execute_reply":"2022-08-06T06:53:04.365922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Bagging reg\nbgreg = BaggingRegressor()\nbgreg.fit(X_train, y_train)\nprint(round(bgreg.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:53:04.370750Z","iopub.execute_input":"2022-08-06T06:53:04.371114Z","iopub.status.idle":"2022-08-06T06:53:04.617537Z","shell.execute_reply.started":"2022-08-06T06:53:04.371082Z","shell.execute_reply":"2022-08-06T06:53:04.616148Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Kernal ridge\nkrid = KernelRidge()\nkrid.fit(X_train, y_train)\nprint(round(krid.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:53:04.619254Z","iopub.execute_input":"2022-08-06T06:53:04.620128Z","iopub.status.idle":"2022-08-06T06:53:04.834456Z","shell.execute_reply.started":"2022-08-06T06:53:04.620076Z","shell.execute_reply":"2022-08-06T06:53:04.832703Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Ridge\nrid = Ridge()\nrid.fit(X_train, y_train)\nprint(round(rid.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:53:04.837056Z","iopub.execute_input":"2022-08-06T06:53:04.839901Z","iopub.status.idle":"2022-08-06T06:53:04.885478Z","shell.execute_reply.started":"2022-08-06T06:53:04.839835Z","shell.execute_reply":"2022-08-06T06:53:04.883991Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Ridge cv\nrigcv = RidgeCV()\nrigcv.fit(X_train, y_train)\nprint(round(rigcv.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:53:04.892362Z","iopub.execute_input":"2022-08-06T06:53:04.896708Z","iopub.status.idle":"2022-08-06T06:53:04.971308Z","shell.execute_reply.started":"2022-08-06T06:53:04.896638Z","shell.execute_reply":"2022-08-06T06:53:04.969334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Elastic net\nelnet = ElasticNet()\nelnet.fit(X_train, y_train)\nprint(round(elnet.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:53:04.979312Z","iopub.execute_input":"2022-08-06T06:53:04.983491Z","iopub.status.idle":"2022-08-06T06:53:05.015896Z","shell.execute_reply.started":"2022-08-06T06:53:04.983395Z","shell.execute_reply":"2022-08-06T06:53:05.014339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Elastic net cv\nelnetcv = ElasticNetCV()\nelnetcv.fit(X_train, y_train)\nprint(round(elnetcv.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:53:05.023341Z","iopub.execute_input":"2022-08-06T06:53:05.027394Z","iopub.status.idle":"2022-08-06T06:53:05.304253Z","shell.execute_reply.started":"2022-08-06T06:53:05.027324Z","shell.execute_reply":"2022-08-06T06:53:05.302137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# SVR\nsvr = SVR()\nsvr.fit(X_train, y_train)\nprint(round(svr.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:53:05.313262Z","iopub.execute_input":"2022-08-06T06:53:05.319123Z","iopub.status.idle":"2022-08-06T06:53:05.893491Z","shell.execute_reply.started":"2022-08-06T06:53:05.319040Z","shell.execute_reply":"2022-08-06T06:53:05.892068Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# LGBMRegressor\nlgbmr = LGBMRegressor()\nlgbmr.fit(X_train, y_train)\nprint(round(lgbmr.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:53:05.894900Z","iopub.execute_input":"2022-08-06T06:53:05.895246Z","iopub.status.idle":"2022-08-06T06:53:06.163726Z","shell.execute_reply.started":"2022-08-06T06:53:05.895217Z","shell.execute_reply":"2022-08-06T06:53:06.162723Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# XGBRegressor\nxgbreg = XGBRegressor()\nxgbreg.fit(X_train, y_train)\nprint(round(xgbreg.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:53:06.168228Z","iopub.execute_input":"2022-08-06T06:53:06.169146Z","iopub.status.idle":"2022-08-06T06:53:06.850783Z","shell.execute_reply.started":"2022-08-06T06:53:06.169103Z","shell.execute_reply":"2022-08-06T06:53:06.849888Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## __Model tuning__","metadata":{}},{"cell_type":"markdown","source":"Lets try these different models with different parameters and find the best algorithm for our solution","metadata":{}},{"cell_type":"code","source":"# define a dictionary for models and their parameters\nmodel_param = {\n    'DicisionsTree': {\n        'model': DecisionTreeRegressor(),\n        'params' : {\n            'random_state': [10,50,100,150,300]\n        }  \n    },\n    'random_forest': {\n        'model': RandomForestRegressor(),\n        'params' : {\n            'n_estimators': [1,5,10,50,100],\n            'random_state': [0,20,50]\n        }\n    },\n    'GradientBoosting' : {\n        'model': GradientBoostingRegressor(),\n        'params': {\n            'n_estimators': [50,100,150], \n            'max_depth': [1,3,5,7], \n            'min_samples_split': [2,3,5],\n            'learning_rate': [0.05]\n        }\n    },\n    'SVR' : {\n        'model' : SVR(),\n        'params' : {\n\n        }\n    },\n    'XGB' : {\n        'model' : XGBRegressor(),\n        'params' : {\n\n        }\n    },\n    'LGB' : {\n        'model' : lgb.LGBMRegressor(),\n        'params' : {\n\n        }\n    },\n    'AdaBoost' : {\n        'model' : AdaBoostRegressor(),\n        'params' : {\n            \n        }\n    },\n    'BaggingRegressor' : {\n        'model' : BaggingRegressor(),\n        'params' : {\n                 \n        }\n    },\n    'KernelRidge' : {\n        'model' : KernelRidge(),\n        'params' : {\n\n        }\n    },\n    'Ridge' : {\n        'model' : Ridge(),\n        'params' : {\n            \n        }\n    },\n    'RidgeCV' : {\n        'model' : RidgeCV(),\n        'params' : {\n            \n        }\n    },\n    'ElasticNet' : {\n        'model' : ElasticNet(),\n        'params' : {\n            \n        }\n    },\n    'ElasticNetCV' : {\n        'model' : ElasticNetCV(),\n        'params' : {\n            \n        }\n    }\n\n}","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:53:53.711388Z","iopub.execute_input":"2022-08-06T06:53:53.712843Z","iopub.status.idle":"2022-08-06T06:53:53.727376Z","shell.execute_reply.started":"2022-08-06T06:53:53.712792Z","shell.execute_reply":"2022-08-06T06:53:53.725872Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#find the best model\nscores = []\nfor model_name, mp in model_param.items():\n  clf = GridSearchCV(mp['model'],mp['params'],cv=3,return_train_score=False)\n  clf.fit(X_train, y_train)\n  scores.append({\n      'model' : model_name,\n      'best_score' : clf.best_score_,\n      'best_params' : clf.best_params_\n  })\ndf = pd.DataFrame(scores,columns=['model','best_score','best_params'])\ndf","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:53:58.089621Z","iopub.execute_input":"2022-08-06T06:53:58.091012Z","iopub.status.idle":"2022-08-06T06:55:43.461928Z","shell.execute_reply.started":"2022-08-06T06:53:58.090956Z","shell.execute_reply":"2022-08-06T06:55:43.460451Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## __Model training and predict__","metadata":{}},{"cell_type":"markdown","source":"So according to above chart we can see lgb is the best scored model with these parameters. So let's use that to train our model.","metadata":{}},{"cell_type":"code","source":"model = lgb.LGBMRegressor(objective='regression',num_leaves=5,learning_rate=0.1, n_estimators=500,max_bin = 55, bagging_fraction = 0.8,bagging_freq = 5, feature_fraction = 0.2319,feature_fraction_seed=9, bagging_seed=9,min_data_in_leaf =6, min_sum_hessian_in_leaf = 11)\nmodel.fit(X_train, y_train)\ny_pred = model.predict(X_test)","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:58:57.087948Z","iopub.execute_input":"2022-08-06T06:58:57.088621Z","iopub.status.idle":"2022-08-06T06:58:57.996943Z","shell.execute_reply.started":"2022-08-06T06:58:57.088568Z","shell.execute_reply":"2022-08-06T06:58:57.995801Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### __Submission__","metadata":{}},{"cell_type":"markdown","source":"Lets do our submission","metadata":{}},{"cell_type":"code","source":"submission = pd.DataFrame({\n    'Id':test_df_copy['Id'],\n    'SalePrice':y_pred\n})\nsubmission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:59:00.744679Z","iopub.execute_input":"2022-08-06T06:59:00.745073Z","iopub.status.idle":"2022-08-06T06:59:00.757971Z","shell.execute_reply.started":"2022-08-06T06:59:00.745041Z","shell.execute_reply":"2022-08-06T06:59:00.757091Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T06:59:02.156416Z","iopub.execute_input":"2022-08-06T06:59:02.156929Z","iopub.status.idle":"2022-08-06T06:59:02.170279Z","shell.execute_reply.started":"2022-08-06T06:59:02.156880Z","shell.execute_reply":"2022-08-06T06:59:02.168965Z"},"trusted":true},"execution_count":null,"outputs":[]}]}