{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Goal\n\n*It is your job to predict the sales price for each house. For each Id in the test set, you must predict the value of the SalePrice variable.* \n\n## Metric\n\n*Submissions are evaluated on Root-Mean-Squared-Error (RMSE) between the logarithm of the predicted value and the logarithm of the observed sales price. (Taking logs means that errors in predicting expensive houses and cheap houses will affect the result equally.)*","metadata":{}},{"cell_type":"code","source":"!pip install sweetviz","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:08:03.627613Z","iopub.execute_input":"2022-08-14T08:08:03.628008Z","iopub.status.idle":"2022-08-14T08:08:20.069467Z","shell.execute_reply.started":"2022-08-14T08:08:03.627977Z","shell.execute_reply":"2022-08-14T08:08:20.068093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"execution":{"iopub.status.busy":"2022-08-14T10:01:19.200382Z","iopub.execute_input":"2022-08-14T10:01:19.200838Z","iopub.status.idle":"2022-08-14T10:01:19.210736Z","shell.execute_reply.started":"2022-08-14T10:01:19.200804Z","shell.execute_reply":"2022-08-14T10:01:19.209493Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## File Descriptions\n\n* train.csv - the training set\n* test.csv - the test set\n* data_description.txt - full description of each column, originally prepared by Dean De Cock but lightly edited to match the column names used here\n* sample_submission.csv - a benchmark submission from a linear regression on year and month of sale, lot square footage, and number of bedrooms\n\n## Data fields\n\n*Here's a brief version of what you'll find in the data description file.*\n\n* SalePrice - the property's sale price in dollars. This is the target variable that you're trying to predict.\n* MSSubClass: The building class\n* MSZoning: The general zoning classification\n* LotFrontage: Linear feet of street connected to property\n* LotArea: Lot size in square feet\n* Street: Type of road access\n* Alley: Type of alley access\n* LotShape: General shape of property\n* LandContour: Flatness of the property\n* Utilities: Type of utilities available\n* LotConfig: Lot configuration\n* LandSlope: Slope of property\n* Neighborhood: Physical locations within Ames city limits\n* Condition1: Proximity to main road or railroad\n* Condition2: Proximity to main road or railroad (if a second is present)\n* BldgType: Type of dwelling\n* HouseStyle: Style of dwelling\n* OverallQual: Overall material and finish quality\n* OverallCond: Overall condition rating\n* YearBuilt: Original construction date\n* YearRemodAdd: Remodel date\n* RoofStyle: Type of roof\n* RoofMatl: Roof material\n* Exterior1st: Exterior covering on house\n* Exterior2nd: Exterior covering on house (if more than one material)\n* MasVnrType: Masonry veneer type\n* MasVnrArea: Masonry veneer area in square feet\n* ExterQual: Exterior material quality\n* ExterCond: Present condition of the material on the exterior\n* Foundation: Type of foundation\n* BsmtQual: Height of the basement\n* BsmtCond: General condition of the basement\n* BsmtExposure: Walkout or garden level basement walls\n* BsmtFinType1: Quality of basement finished area\n* BsmtFinSF1: Type 1 finished square feet\n* BsmtFinType2: Quality of second finished area (if present)\n* BsmtFinSF2: Type 2 finished square feet\n* BsmtUnfSF: Unfinished square feet of basement area\n* TotalBsmtSF: Total square feet of basement area\n* Heating: Type of heating\n* HeatingQC: Heating quality and condition\n* CentralAir: Central air conditioning\n* Electrical: Electrical system\n* 1stFlrSF: First Floor square feet\n* 2ndFlrSF: Second floor square feet\n* LowQualFinSF: Low quality finished square feet (all floors)\n* GrLivArea: Above grade (ground) living area square feet\n* BsmtFullBath: Basement full bathrooms\n* BsmtHalfBath: Basement half bathrooms\n* FullBath: Full bathrooms above grade\n* HalfBath: Half baths above grade\n* Bedroom: Number of bedrooms above basement level\n* Kitchen: Number of kitchens\n* KitchenQual: Kitchen quality\n* TotRmsAbvGrd: Total rooms above grade (does not include bathrooms)\n* Functional: Home functionality rating\n* Fireplaces: Number of fireplaces\n* FireplaceQu: Fireplace quality\n* GarageType: Garage location\n* GarageYrBlt: Year garage was built\n* GarageFinish: Interior finish of the garage\n* GarageCars: Size of garage in car capacity\n* GarageArea: Size of garage in square feet\n* GarageQual: Garage quality\n* GarageCond: Garage condition\n* PavedDrive: Paved driveway\n* WoodDeckSF: Wood deck area in square feet\n* OpenPorchSF: Open porch area in square feet\n* EnclosedPorch: Enclosed porch area in square feet\n* 3SsnPorch: Three season porch area in square feet\n* ScreenPorch: Screen porch area in square feet\n* PoolArea: Pool area in square feet\n* PoolQC: Pool quality\n* Fence: Fence quality\n* MiscFeature: Miscellaneous feature not covered in other categories\n* MiscVal: $Value of miscellaneous feature\n* MoSold: Month Sold\n* YrSold: Year Sold\n* SaleType: Type of sale\n* SaleCondition: Condition of sale","metadata":{}},{"cell_type":"markdown","source":"## Let's Import the Libraries ","metadata":{}},{"cell_type":"code","source":"import pandas as pd \nimport seaborn as sns\nimport matplotlib.pyplot as plt\nsns.set_style('darkgrid')\nfrom sklearn.metrics import mean_absolute_error\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.preprocessing import OneHotEncoder\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.model_selection import cross_val_score, KFold\nfrom sklearn.metrics import mean_squared_error\nimport sweetviz as sv\n\nimport xgboost as xgb","metadata":{"execution":{"iopub.status.busy":"2022-08-14T10:01:21.154853Z","iopub.execute_input":"2022-08-14T10:01:21.155341Z","iopub.status.idle":"2022-08-14T10:01:21.164027Z","shell.execute_reply.started":"2022-08-14T10:01:21.155296Z","shell.execute_reply":"2022-08-14T10:01:21.162027Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Let's Load the Data","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/home-data-for-ml-course/train.csv')\ntest = pd.read_csv('/kaggle/input/home-data-for-ml-course/test.csv')\nsample_sub = pd.read_csv('/kaggle/input/home-data-for-ml-course/sample_submission.csv')\n\n# Let's print the head rows of training set\n\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T10:01:23.944470Z","iopub.execute_input":"2022-08-14T10:01:23.945203Z","iopub.status.idle":"2022-08-14T10:01:24.013037Z","shell.execute_reply.started":"2022-08-14T10:01:23.945165Z","shell.execute_reply":"2022-08-14T10:01:24.011789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's check dimesions of our train and test set\nprint(f\"Dimensions of training set is : {train.shape}\")\nprint(f\"Dimensions of testing set is : {test.shape}\")","metadata":{"execution":{"iopub.status.busy":"2022-08-14T10:01:24.723154Z","iopub.execute_input":"2022-08-14T10:01:24.723921Z","iopub.status.idle":"2022-08-14T10:01:24.731233Z","shell.execute_reply.started":"2022-08-14T10:01:24.723872Z","shell.execute_reply":"2022-08-14T10:01:24.729511Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can clearly see that there are 81 columns in train and 80 columns in test set, Out of which we don't need Id column in our train set for the training purpose as well as we don't need for test set","metadata":{}},{"cell_type":"code","source":"train.drop(columns = ['Id'], inplace = True)\ntest.drop(columns = ['Id'], inplace = True)\n\ntrain.shape, test.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-14T10:01:25.941039Z","iopub.execute_input":"2022-08-14T10:01:25.942488Z","iopub.status.idle":"2022-08-14T10:01:25.957481Z","shell.execute_reply.started":"2022-08-14T10:01:25.942421Z","shell.execute_reply":"2022-08-14T10:01:25.955679Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## EDA (Exploratory Data Analysis)","metadata":{}},{"cell_type":"code","source":"train.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T10:01:27.422777Z","iopub.execute_input":"2022-08-14T10:01:27.423595Z","iopub.status.idle":"2022-08-14T10:01:27.451188Z","shell.execute_reply.started":"2022-08-14T10:01:27.423533Z","shell.execute_reply":"2022-08-14T10:01:27.449707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's use sweetviz library great functionality to show us some good visuals and give us pretty useful info about our training data\nviz = sv.analyze(train)\nviz.show_notebook()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T10:01:28.414059Z","iopub.execute_input":"2022-08-14T10:01:28.415291Z","iopub.status.idle":"2022-08-14T10:02:21.510589Z","shell.execute_reply.started":"2022-08-14T10:01:28.415244Z","shell.execute_reply":"2022-08-14T10:02:21.508194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, we can clearly see from the above visuals from sweetviz libaray function and also test from .info() functions that there are alot of missing values in some columns, in total we have 79 columns in which 54 are categorical, 24 are numericals and one is text column.","metadata":{}},{"cell_type":"code","source":"train.isna().sum().sort_values(ascending = False).head(15)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T10:02:21.513699Z","iopub.execute_input":"2022-08-14T10:02:21.514299Z","iopub.status.idle":"2022-08-14T10:02:21.530053Z","shell.execute_reply.started":"2022-08-14T10:02:21.514263Z","shell.execute_reply":"2022-08-14T10:02:21.529030Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's drop top 6 columns with most null values and then fill the missing values in other columns\ntrain.drop(columns = ['PoolQC', 'MiscFeature', 'Alley', 'Fence', 'FireplaceQu', 'LotFrontage'], inplace = True)\ntest.drop(columns = ['PoolQC', 'MiscFeature', 'Alley', 'Fence', 'FireplaceQu', 'LotFrontage'], inplace = True)\n\n# Let's print dimensions of train and test set\ntrain.shape, test.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-14T10:02:21.531809Z","iopub.execute_input":"2022-08-14T10:02:21.532797Z","iopub.status.idle":"2022-08-14T10:02:21.551265Z","shell.execute_reply.started":"2022-08-14T10:02:21.532725Z","shell.execute_reply":"2022-08-14T10:02:21.549978Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, let's just try to fill the missing values for other columns","metadata":{}},{"cell_type":"code","source":"train['MSZoning'] = train['MSZoning'].fillna(value='RL')\ntrain['Exterior1st'] = train['Exterior1st'].fillna(value='VinylSd')\ntrain['Exterior2nd'] = train['Exterior2nd'].fillna(value='VinylSd')\ntrain['MasVnrType'] = train['MasVnrType'].fillna(value='None')\ntrain['MasVnrArea'] = train['MasVnrArea'].fillna(value=0)\ntrain[['BsmtQual','BsmtCond','BsmtExposure','BsmtFinType1','BsmtFinType2']] = train[['BsmtQual','BsmtCond','BsmtExposure','BsmtFinType1','BsmtFinType2']].fillna(value='None')\ntrain[['BsmtFinSF1','BsmtFinSF2','BsmtUnfSF','TotalBsmtSF','BsmtFullBath','BsmtHalfBath']] = train[['BsmtFinSF1','BsmtFinSF2','BsmtUnfSF','TotalBsmtSF','BsmtFullBath','BsmtHalfBath']].fillna(value=0)\ntrain['Electrical'] = train['Electrical'].fillna(value='SBrkr')\ntrain['KitchenQual'] = train['KitchenQual'].fillna(value='TA')\ntrain['Functional'] = train['Functional'].fillna(value='Typ')\ntrain[['GarageType','GarageFinish','GarageQual','GarageCond']] = train[['GarageType','GarageFinish','GarageQual','GarageCond']].fillna(value='None')\ntrain[['GarageCars','GarageArea']] = train[['GarageCars','GarageArea']].fillna(value=0)\ntrain['GarageYrBlt'] = train['GarageYrBlt'].fillna(value=1900)\n\ntrain['SaleType'] = train['SaleType'].fillna(value='WD')\n\n# Let's do the same with test data\ntest['MSZoning'] = test['MSZoning'].fillna(value='RL')\ntest['Exterior1st'] = test['Exterior1st'].fillna(value='VinylSd')\ntest['Exterior2nd'] = test['Exterior2nd'].fillna(value='VinylSd')\ntest['MasVnrType'] = test['MasVnrType'].fillna(value='None')\ntest['MasVnrArea'] = test['MasVnrArea'].fillna(value=0)\ntest[['BsmtQual','BsmtCond','BsmtExposure','BsmtFinType1','BsmtFinType2']] = test[['BsmtQual','BsmtCond','BsmtExposure','BsmtFinType1','BsmtFinType2']].fillna(value='None')\ntest[['BsmtFinSF1','BsmtFinSF2','BsmtUnfSF','TotalBsmtSF','BsmtFullBath','BsmtHalfBath']] = test[['BsmtFinSF1','BsmtFinSF2','BsmtUnfSF','TotalBsmtSF','BsmtFullBath','BsmtHalfBath']].fillna(value=0)\ntest['Electrical'] = test['Electrical'].fillna(value='SBrkr')\ntest['KitchenQual'] = test['KitchenQual'].fillna(value='TA')\ntest['Functional'] = test['Functional'].fillna(value='Typ')\ntest[['GarageType','GarageFinish','GarageQual','GarageCond']] = test[['GarageType','GarageFinish','GarageQual','GarageCond']].fillna(value='None')\ntest[['GarageCars','GarageArea']] = test[['GarageCars','GarageArea']].fillna(value=0)\ntest['GarageYrBlt'] = test['GarageYrBlt'].fillna(value=1900)\n\ntest['SaleType'] = test['SaleType'].fillna(value='WD')","metadata":{"execution":{"iopub.status.busy":"2022-08-14T10:02:21.554527Z","iopub.execute_input":"2022-08-14T10:02:21.555230Z","iopub.status.idle":"2022-08-14T10:02:21.603779Z","shell.execute_reply.started":"2022-08-14T10:02:21.555163Z","shell.execute_reply":"2022-08-14T10:02:21.602707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# let's check train data again\ntrain.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T10:02:21.605463Z","iopub.execute_input":"2022-08-14T10:02:21.606138Z","iopub.status.idle":"2022-08-14T10:02:21.629527Z","shell.execute_reply.started":"2022-08-14T10:02:21.606099Z","shell.execute_reply":"2022-08-14T10:02:21.628653Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"all the null values are successfully either removed or filled","metadata":{}},{"cell_type":"markdown","source":"## LabelEncoder\n\nNow, let's use label encoder to encode categorical variables into numerical, so that we can train our algorithms on it successfully.","metadata":{}},{"cell_type":"code","source":"# let's first extract categorical variables from training set\ncategorical_variables = list(train.select_dtypes(exclude = ['int64', 'float64']).columns)\n\nfor var in categorical_variables:\n    label_encod = LabelEncoder()\n    arr = np.concatenate((train[var], test[var])).astype(str)\n    label_encod.fit(arr)\n    train[var] = label_encod.transform(train[var].astype(str))\n    test[var]  = label_encod.transform(test[var].astype(str))","metadata":{"execution":{"iopub.status.busy":"2022-08-14T10:02:21.630965Z","iopub.execute_input":"2022-08-14T10:02:21.631885Z","iopub.status.idle":"2022-08-14T10:02:21.755985Z","shell.execute_reply.started":"2022-08-14T10:02:21.631847Z","shell.execute_reply":"2022-08-14T10:02:21.755045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Scatterplot\n\nLet's plot scatterplot of predictor variables to see the correlation among them","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize = (22, 16))\ncorrela = train.corr()\nsns.heatmap(correla)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T10:02:21.757493Z","iopub.execute_input":"2022-08-14T10:02:21.758025Z","iopub.status.idle":"2022-08-14T10:02:24.413703Z","shell.execute_reply.started":"2022-08-14T10:02:21.757992Z","shell.execute_reply":"2022-08-14T10:02:24.412466Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Due, to huge number of variables the values cannot be seen clearly but still with the color intensity one can understand the correlation among variables if one look closely. But again due to great number of variables this heatmap is just not useful","metadata":{}},{"cell_type":"markdown","source":"## Standard Scaler\n\nNow, we will use standard scaler library from  Scikit learn to apply feature scalling to both train and test set.","metadata":{}},{"cell_type":"code","source":"# Let's first assign SalePrice to Y as a target variable and all the other variables to X as predictor variables\ntrain['SalePrice'] = np.log1p(train['SalePrice'])\nY = train[['SalePrice']]\nX = train.drop(['SalePrice'], axis = 1)\n\nfeatures = list(X.columns)\nscaler = StandardScaler()\nX[features] = scaler.fit_transform(X[features])\ntest[features] = scaler.fit_transform(test[features])","metadata":{"execution":{"iopub.status.busy":"2022-08-14T10:17:33.579344Z","iopub.execute_input":"2022-08-14T10:17:33.579839Z","iopub.status.idle":"2022-08-14T10:17:33.630111Z","shell.execute_reply.started":"2022-08-14T10:17:33.579802Z","shell.execute_reply":"2022-08-14T10:17:33.628902Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model fitting and Testing\n\nNow, we have the data ready for applying XGBregressor on it and below we will first fit the model and than will check it's accuracy and then will apply predict function of trained model on test and do the submission.","metadata":{}},{"cell_type":"code","source":"# Initialization of xgbregressor\nmodel = xgb.XGBRegressor()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T10:17:35.270590Z","iopub.execute_input":"2022-08-14T10:17:35.271550Z","iopub.status.idle":"2022-08-14T10:17:35.278690Z","shell.execute_reply.started":"2022-08-14T10:17:35.271505Z","shell.execute_reply":"2022-08-14T10:17:35.277043Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# fitting of the model on training set\nmodel.fit(X, Y, verbose=True)\nmodel","metadata":{"execution":{"iopub.status.busy":"2022-08-14T10:17:37.916111Z","iopub.execute_input":"2022-08-14T10:17:37.916618Z","iopub.status.idle":"2022-08-14T10:17:38.628776Z","shell.execute_reply.started":"2022-08-14T10:17:37.916578Z","shell.execute_reply":"2022-08-14T10:17:38.627776Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's check model accuracy on training data\nprint(model.score(X, Y))","metadata":{"execution":{"iopub.status.busy":"2022-08-14T10:17:39.292901Z","iopub.execute_input":"2022-08-14T10:17:39.293359Z","iopub.status.idle":"2022-08-14T10:17:39.318668Z","shell.execute_reply.started":"2022-08-14T10:17:39.293325Z","shell.execute_reply":"2022-08-14T10:17:39.317695Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's check cross vaildation accuracy on training data\nscores = cross_val_score(model, X, Y,cv=10)\nprint(\"Mean cross-validation score: %.2f\" % scores.mean())","metadata":{"execution":{"iopub.status.busy":"2022-08-14T10:17:40.637617Z","iopub.execute_input":"2022-08-14T10:17:40.638063Z","iopub.status.idle":"2022-08-14T10:17:47.970732Z","shell.execute_reply.started":"2022-08-14T10:17:40.638030Z","shell.execute_reply":"2022-08-14T10:17:47.969391Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can clearly see that model accuracy on training data is almost 100%, but average score of cross validation score with cv = 10 is 88% which shows that our model is doing overfitting and we need to do something otherwise model performance on test data will not be good.","metadata":{}},{"cell_type":"code","source":"model1 = xgb.XGBRegressor(base_score=0.5, booster='gbtree', callbacks=None,\n             colsample_bylevel=1, colsample_bynode=1, colsample_bytree=1,\n             learning_rate=0.32, max_bin=256, max_cat_to_onehot=4,\n             max_delta_step=0, max_depth=4, max_leaves=0, min_child_weight=1,\n              monotone_constraints='()', n_estimators=20, n_jobs=0,\n             num_parallel_tree=1, predictor='auto', random_state=0, reg_alpha=0.8,\n             reg_lambda=7)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T10:13:04.772918Z","iopub.execute_input":"2022-08-14T10:13:04.773359Z","iopub.status.idle":"2022-08-14T10:13:04.781294Z","shell.execute_reply.started":"2022-08-14T10:13:04.773325Z","shell.execute_reply":"2022-08-14T10:13:04.780064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model1.fit(X, Y, verbose=True)\nmodel1","metadata":{"execution":{"iopub.status.busy":"2022-08-14T10:13:05.486827Z","iopub.execute_input":"2022-08-14T10:13:05.487748Z","iopub.status.idle":"2022-08-14T10:13:05.636329Z","shell.execute_reply.started":"2022-08-14T10:13:05.487707Z","shell.execute_reply":"2022-08-14T10:13:05.635133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's check model accuracy on training data\nprint(model1.score(X, Y))","metadata":{"execution":{"iopub.status.busy":"2022-08-14T10:13:06.500855Z","iopub.execute_input":"2022-08-14T10:13:06.501585Z","iopub.status.idle":"2022-08-14T10:13:06.520438Z","shell.execute_reply.started":"2022-08-14T10:13:06.501528Z","shell.execute_reply":"2022-08-14T10:13:06.519464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's check cross vaildation accuracy on training data\nscores = cross_val_score(model1, X, Y,cv=10)\nprint(\"Mean cross-validation score: %.2f\" % scores.mean())","metadata":{"execution":{"iopub.status.busy":"2022-08-14T10:13:08.616233Z","iopub.execute_input":"2022-08-14T10:13:08.617609Z","iopub.status.idle":"2022-08-14T10:13:10.036228Z","shell.execute_reply.started":"2022-08-14T10:13:08.617556Z","shell.execute_reply":"2022-08-14T10:13:10.035316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's use model.predict function to predict SalePrice for test data\n\npredictions = model.predict(test)\nsample_sub.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T10:17:55.645666Z","iopub.execute_input":"2022-08-14T10:17:55.646071Z","iopub.status.idle":"2022-08-14T10:17:55.666585Z","shell.execute_reply.started":"2022-08-14T10:17:55.646040Z","shell.execute_reply":"2022-08-14T10:17:55.665181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_sub.SalePrice = predictions\nsample_sub.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T10:17:56.467510Z","iopub.execute_input":"2022-08-14T10:17:56.467916Z","iopub.status.idle":"2022-08-14T10:17:56.480986Z","shell.execute_reply.started":"2022-08-14T10:17:56.467886Z","shell.execute_reply":"2022-08-14T10:17:56.479531Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_sub.to_csv('submission.csv', index = False)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T10:18:08.403084Z","iopub.execute_input":"2022-08-14T10:18:08.403538Z","iopub.status.idle":"2022-08-14T10:18:08.414534Z","shell.execute_reply.started":"2022-08-14T10:18:08.403504Z","shell.execute_reply":"2022-08-14T10:18:08.413680Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}