{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## *Housing Prices Competition for Kaggle Learn Users*","metadata":{}},{"cell_type":"markdown","source":"***Upvote my Notebook :)***","metadata":{}},{"cell_type":"markdown","source":"# Important Libraries:","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.model_selection import train_test_split, RandomizedSearchCV\nfrom sklearn.ensemble import RandomForestRegressor,GradientBoostingRegressor\nimport xgboost as xgb\nfrom sklearn.metrics import mean_absolute_error\nimport warnings\nwarnings.filterwarnings(\"ignore\")","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2022-07-24T08:59:41.153117Z","iopub.execute_input":"2022-07-24T08:59:41.154081Z","iopub.status.idle":"2022-07-24T08:59:42.677052Z","shell.execute_reply.started":"2022-07-24T08:59:41.153928Z","shell.execute_reply":"2022-07-24T08:59:42.675859Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# First we analyze the train and test Data.","metadata":{}},{"cell_type":"code","source":"train_data = pd.read_csv('../input/home-data-for-ml-course/train.csv')\ntest_data = pd.read_csv('../input/home-data-for-ml-course/test.csv')","metadata":{"execution":{"iopub.status.busy":"2022-07-24T09:00:16.766067Z","iopub.execute_input":"2022-07-24T09:00:16.766511Z","iopub.status.idle":"2022-07-24T09:00:16.855508Z","shell.execute_reply.started":"2022-07-24T09:00:16.766469Z","shell.execute_reply":"2022-07-24T09:00:16.854396Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* **We check that how many columns are there and what type of columns are there to gain some knowledge from our newly loaded data.**\n* **Then we also get to know that how many rows on particular column are null or non-null**","metadata":{}},{"cell_type":"code","source":"train_data.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T19:27:14.539711Z","iopub.execute_input":"2022-07-21T19:27:14.540454Z","iopub.status.idle":"2022-07-21T19:27:14.566470Z","shell.execute_reply.started":"2022-07-21T19:27:14.540401Z","shell.execute_reply":"2022-07-21T19:27:14.565263Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* **So, to Make it easier to analyze the data. we split the train_data (Train.csv Data) into two parts:**\n* **One is int+float type parts and One is Object type parts**\n* **By part means I have spilts the table's columns into two parts.**","metadata":{}},{"cell_type":"code","source":"num_val = train_data.select_dtypes(exclude=['object']).copy()\ncat_val = train_data.select_dtypes(include=['object']).copy()","metadata":{"execution":{"iopub.status.busy":"2022-07-24T09:00:20.993632Z","iopub.execute_input":"2022-07-24T09:00:20.994638Z","iopub.status.idle":"2022-07-24T09:00:21.012590Z","shell.execute_reply.started":"2022-07-24T09:00:20.994583Z","shell.execute_reply":"2022-07-24T09:00:21.011527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* **Checking the num_val which contain numerical values,**","metadata":{}},{"cell_type":"code","source":"num_val.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-24T09:00:23.749889Z","iopub.execute_input":"2022-07-24T09:00:23.750580Z","iopub.status.idle":"2022-07-24T09:00:23.777907Z","shell.execute_reply.started":"2022-07-24T09:00:23.750543Z","shell.execute_reply":"2022-07-24T09:00:23.776957Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* **Trying to Know that how much data from num_val contain Nan values and which column contain the most the upto so on by visualizing we get a clear picture in mind that there are few columns which contain Nan.**\n* **Now we apply fillna function which is used to fill those cells of particular column with mean,median,mode or whatever value we choose to fill.**","metadata":{}},{"cell_type":"markdown","source":"*** So Column : LotFrontage, MasVnrArea, GarageYrBlt have Nan values**********","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(12,6))\nplt.title('Null Values Count')\nnum_val.isnull().sum().plot(kind='bar',legend=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-24T09:00:30.255631Z","iopub.execute_input":"2022-07-24T09:00:30.256016Z","iopub.status.idle":"2022-07-24T09:00:30.883434Z","shell.execute_reply.started":"2022-07-24T09:00:30.255964Z","shell.execute_reply":"2022-07-24T09:00:30.882493Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_val['LotFrontage'].fillna(num_val['LotFrontage'].median(),inplace=True)\nnum_val['MasVnrArea'].fillna(num_val['MasVnrArea'].median(),inplace=True)\nnum_val['GarageYrBlt'].fillna(num_val['GarageYrBlt'].mode()[0],inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-24T09:00:37.929655Z","iopub.execute_input":"2022-07-24T09:00:37.930029Z","iopub.status.idle":"2022-07-24T09:00:37.942638Z","shell.execute_reply.started":"2022-07-24T09:00:37.929990Z","shell.execute_reply":"2022-07-24T09:00:37.941561Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* **Plotting Histogram to check skewness from the features**\n* **Skewed columns : LowQualFinSF, 3SsnPorch, PoolArea, MiscVal,LotArea**\n* **We apply log transformation in those columns which contain continuous values. we can not apply log transformation on non-continuous columns because these columns will be filled with infinity or some Nan values which we don't want.**\n* **So, we apply log transformation to LotArea column which is the continuous column and we will drop other 4 columns**","metadata":{}},{"cell_type":"code","source":"fig = plt.figure(figsize=(22,24))\nfor ind,col in enumerate(num_val):\n    plt.subplot(6,7,ind+1)\n    sns.histplot(data=num_val.loc[:,col].dropna(),kde=False,color='red')\n\nfig.tight_layout(pad=1.5)","metadata":{"execution":{"iopub.status.busy":"2022-07-24T09:00:40.630761Z","iopub.execute_input":"2022-07-24T09:00:40.631463Z","iopub.status.idle":"2022-07-24T09:00:47.959522Z","shell.execute_reply.started":"2022-07-24T09:00:40.631421Z","shell.execute_reply":"2022-07-24T09:00:47.958555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Log Transformation\nnum_val['LotArea'] = np.log(num_val['LotArea'])","metadata":{"execution":{"iopub.status.busy":"2022-07-24T09:00:52.367870Z","iopub.execute_input":"2022-07-24T09:00:52.368511Z","iopub.status.idle":"2022-07-24T09:00:52.375466Z","shell.execute_reply.started":"2022-07-24T09:00:52.368471Z","shell.execute_reply":"2022-07-24T09:00:52.374054Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"skewed_cols = ['LowQualFinSF','3SsnPorch','PoolArea','MiscVal']\nnum_val.drop(skewed_cols,axis=1,inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-24T09:00:55.752946Z","iopub.execute_input":"2022-07-24T09:00:55.753829Z","iopub.status.idle":"2022-07-24T09:00:55.760630Z","shell.execute_reply.started":"2022-07-24T09:00:55.753778Z","shell.execute_reply":"2022-07-24T09:00:55.759640Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Now we check Variance of multiple columns because High Variance can make our model overfit. ","metadata":{}},{"cell_type":"markdown","source":"* **So No column contain too high variance**","metadata":{}},{"cell_type":"code","source":"num_val.var().sort_values(ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-24T09:00:58.829073Z","iopub.execute_input":"2022-07-24T09:00:58.830055Z","iopub.status.idle":"2022-07-24T09:00:58.844002Z","shell.execute_reply.started":"2022-07-24T09:00:58.830007Z","shell.execute_reply":"2022-07-24T09:00:58.843052Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* **Our Data can contain outliers, so to tackle this kind of problem we first need to calculate some statistical forms.**\n* **we calculate firstQuartile,thirdQuartile and IQR the check the outliers**","metadata":{}},{"cell_type":"code","source":"# Finding Outliers \nfor col in num_val.columns:\n    \n    first_quartile = num_val[col].quantile(0.25) \n    third_quartile = num_val[col].quantile(0.75)\n\n    IQR = third_quartile - first_quartile\n    out = third_quartile + 3*IQR \n    num_val.drop(num_val[num_val[col] > out].index,axis=0,inplace=True)\n","metadata":{"execution":{"iopub.status.busy":"2022-07-24T09:01:03.439563Z","iopub.execute_input":"2022-07-24T09:01:03.439911Z","iopub.status.idle":"2022-07-24T09:01:03.587471Z","shell.execute_reply.started":"2022-07-24T09:01:03.439880Z","shell.execute_reply":"2022-07-24T09:01:03.586292Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize=(22,24))\nfor ind,col in enumerate(num_val):\n    plt.subplot(6,7,ind+1)\n    sns.boxplot(x=num_val.loc[:,col],data=num_val)","metadata":{"execution":{"iopub.status.busy":"2022-07-24T09:01:33.482618Z","iopub.execute_input":"2022-07-24T09:01:33.483029Z","iopub.status.idle":"2022-07-24T09:01:36.453282Z","shell.execute_reply.started":"2022-07-24T09:01:33.482971Z","shell.execute_reply":"2022-07-24T09:01:36.449022Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* **Check the Correlation between Columns of our data.**\n* **Those columns have too high correlation like more than 0.75, then we will drop those columns.**","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(22,24))\nsns.heatmap(num_val.corr() > 0.75,annot=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-24T09:01:41.503435Z","iopub.execute_input":"2022-07-24T09:01:41.504228Z","iopub.status.idle":"2022-07-24T09:01:44.957113Z","shell.execute_reply.started":"2022-07-24T09:01:41.504190Z","shell.execute_reply":"2022-07-24T09:01:44.955555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_val.drop(['GarageCars','Id'],axis=1,inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-24T09:01:48.732650Z","iopub.execute_input":"2022-07-24T09:01:48.733020Z","iopub.status.idle":"2022-07-24T09:01:48.739187Z","shell.execute_reply.started":"2022-07-24T09:01:48.732968Z","shell.execute_reply":"2022-07-24T09:01:48.738109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Now we Check our Categorical Data","metadata":{}},{"cell_type":"markdown","source":"* **first we check that which columns contain Nan values.**\n* **Some Columns have high Nan values which can not be improve by filling there Nan values. So, we will drop those few columns**\n* **We Drop Columns : Alley, PoolQC, MiscFeature, Fence**","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(22,24))\ncat_val.isnull().sum().plot(kind='bar',legend=True,color='forestgreen')","metadata":{"execution":{"iopub.status.busy":"2022-07-24T09:02:03.722593Z","iopub.execute_input":"2022-07-24T09:02:03.723088Z","iopub.status.idle":"2022-07-24T09:02:04.566721Z","shell.execute_reply.started":"2022-07-24T09:02:03.723033Z","shell.execute_reply":"2022-07-24T09:02:04.565739Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"drop_col = ['Alley','PoolQC','MiscFeature','Fence']\ncat_val.drop(drop_col,axis=1,inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-24T09:02:08.446560Z","iopub.execute_input":"2022-07-24T09:02:08.447040Z","iopub.status.idle":"2022-07-24T09:02:08.454800Z","shell.execute_reply.started":"2022-07-24T09:02:08.446985Z","shell.execute_reply":"2022-07-24T09:02:08.453794Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**We filled other columns with Mode, we can not use Mean or median on these categorical data and Mode is the best fit for categorical columns also.**","metadata":{}},{"cell_type":"code","source":"cat_val['MasVnrType'].fillna(cat_val['MasVnrType'].mode()[0],inplace=True)\ncat_val['BsmtQual'].fillna(cat_val['BsmtQual'].mode()[0],inplace=True)\ncat_val['BsmtCond'].fillna(cat_val['BsmtCond'].mode()[0],inplace=True)\ncat_val['BsmtExposure'].fillna(cat_val['BsmtExposure'].mode()[0],inplace=True)\ncat_val['BsmtFinType1'].fillna(cat_val['BsmtFinType1'].mode()[0],inplace=True)\ncat_val['BsmtFinType2'].fillna(cat_val['BsmtFinType2'].mode()[0],inplace=True)\ncat_val['FireplaceQu'].fillna(cat_val['FireplaceQu'].mode()[0],inplace=True)\ncat_val['GarageType'].fillna(cat_val['GarageType'].mode()[0],inplace=True)\ncat_val['GarageFinish'].fillna(cat_val['GarageFinish'].mode()[0],inplace=True)\ncat_val['GarageQual'].fillna(cat_val['GarageQual'].mode()[0],inplace=True)\ncat_val['GarageCond'].fillna(cat_val['GarageCond'].mode()[0],inplace=True)\ncat_val['Electrical'].fillna(cat_val['Electrical'].mode()[0],inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-24T09:02:10.495226Z","iopub.execute_input":"2022-07-24T09:02:10.496227Z","iopub.status.idle":"2022-07-24T09:02:10.521344Z","shell.execute_reply.started":"2022-07-24T09:02:10.496180Z","shell.execute_reply":"2022-07-24T09:02:10.520370Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Now we analyzed Our Data and Know we apply all the above methods that we done by splitting it. we know apply those functions in train_data and test_data**","metadata":{}},{"cell_type":"code","source":"columns = ['MSSubClass', 'LotFrontage', 'LotArea', 'OverallQual', 'OverallCond',\n       'YearBuilt', 'YearRemodAdd', 'MasVnrArea', 'BsmtFinSF1', 'BsmtFinSF2',\n       'BsmtUnfSF', 'TotalBsmtSF', '1stFlrSF', '2ndFlrSF', 'GrLivArea',\n       'BsmtFullBath', 'BsmtHalfBath', 'FullBath', 'HalfBath', 'BedroomAbvGr',\n       'KitchenAbvGr', 'TotRmsAbvGrd', 'Fireplaces',\n       'GarageArea', 'WoodDeckSF', 'OpenPorchSF', 'EnclosedPorch',\n       'ScreenPorch', 'MoSold', 'YrSold']","metadata":{"execution":{"iopub.status.busy":"2022-07-24T09:02:13.333259Z","iopub.execute_input":"2022-07-24T09:02:13.334209Z","iopub.status.idle":"2022-07-24T09:02:13.340541Z","shell.execute_reply.started":"2022-07-24T09:02:13.334167Z","shell.execute_reply":"2022-07-24T09:02:13.339225Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# In this cell we applied all those above Functions to make our train_data cleaner and insight full","metadata":{}},{"cell_type":"code","source":"# Numerical :\n\ntrain_data['LotFrontage'].fillna(train_data['LotFrontage'].median(),inplace=True)\ntrain_data['MasVnrArea'].fillna(train_data['MasVnrArea'].median(),inplace=True)\n\n# Log Transformation: Applied on High Variance columns\n\n\ntrain_data['LotArea'] = np.log(train_data['LotArea'])\n\n# Dropped Skewed Columns\n\nskewed_cols = ['LowQualFinSF','3SsnPorch','PoolArea','MiscVal','GarageYrBlt']\ntrain_data.drop(skewed_cols,axis=1,inplace=True)\n\n# Finding Outliers \nfor col in columns:\n    \n    first_quartile = train_data[col].quantile(0.25) \n    third_quartile = train_data[col].quantile(0.75)\n\n    IQR = third_quartile - first_quartile\n    out = third_quartile + 3*IQR \n    train_data.drop(train_data[train_data[col] > out].index,axis=0,inplace=True)\n\n\n\ntrain_data.drop(['Id','GarageCars'],axis=1,inplace=True)\n\n# Categorical : \n\n\ndrop_col = ['Alley','PoolQC','MiscFeature','Fence']\ntrain_data.drop(drop_col,axis=1,inplace=True)\n\ntrain_data['MasVnrType'].fillna(train_data['MasVnrType'].mode()[0],inplace=True)\ntrain_data['BsmtQual'].fillna(train_data['BsmtQual'].mode()[0],inplace=True)\ntrain_data['BsmtCond'].fillna(train_data['BsmtCond'].mode()[0],inplace=True)\ntrain_data['BsmtExposure'].fillna(train_data['BsmtExposure'].mode()[0],inplace=True)\ntrain_data['BsmtFinType1'].fillna(train_data['BsmtFinType1'].mode()[0],inplace=True)\ntrain_data['BsmtFinType2'].fillna(train_data['BsmtFinType2'].mode()[0],inplace=True)\ntrain_data['FireplaceQu'].fillna(train_data['FireplaceQu'].mode()[0],inplace=True)\ntrain_data['GarageType'].fillna(train_data['GarageType'].mode()[0],inplace=True)\ntrain_data['GarageFinish'].fillna(train_data['GarageFinish'].mode()[0],inplace=True)\ntrain_data['GarageQual'].fillna(train_data['GarageQual'].mode()[0],inplace=True)\ntrain_data['GarageCond'].fillna(train_data['GarageCond'].mode()[0],inplace=True)\ntrain_data['Electrical'].fillna(train_data['Electrical'].mode()[0],inplace=True)\n","metadata":{"execution":{"iopub.status.busy":"2022-07-21T19:29:06.406926Z","iopub.execute_input":"2022-07-21T19:29:06.407926Z","iopub.status.idle":"2022-07-21T19:29:06.514603Z","shell.execute_reply.started":"2022-07-21T19:29:06.407874Z","shell.execute_reply":"2022-07-21T19:29:06.513620Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* **We know apply our last function of pre processing, which is LabelEncoder**\n* **LabelEncoder Function/Method will encode the categorical data into Numerical**\n* **We applied Encoding because Our Machine Learning Model can not be applied on string data, they can only be applicable on numerical data**","metadata":{}},{"cell_type":"code","source":"cat_columns = ['MSZoning','LotShape', 'LandContour', 'Utilities',\n       'LotConfig', 'LandSlope', 'Neighborhood', 'Condition1',\n       'BldgType', 'HouseStyle', 'RoofStyle','Exterior1st',\n       'Exterior2nd', 'MasVnrType', 'ExterQual', 'ExterCond', 'Foundation',\n       'BsmtQual', 'BsmtCond', 'BsmtExposure', 'BsmtFinType1', 'BsmtFinType2',\n       'HeatingQC', 'CentralAir','KitchenQual','Functional', 'FireplaceQu', \n        'GarageType', 'GarageFinish', 'GarageQual','Condition2',\n       'GarageCond', 'PavedDrive', 'Electrical','Street','RoofMatl', 'Heating',\n       'SaleType', 'SaleCondition']\n","metadata":{"execution":{"iopub.status.busy":"2022-07-21T19:29:10.202056Z","iopub.execute_input":"2022-07-21T19:29:10.202399Z","iopub.status.idle":"2022-07-21T19:29:10.208648Z","shell.execute_reply.started":"2022-07-21T19:29:10.202369Z","shell.execute_reply":"2022-07-21T19:29:10.207724Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"le = LabelEncoder()\nfor col in cat_columns:\n    train_data[col] = le.fit_transform(train_data[col])","metadata":{"execution":{"iopub.status.busy":"2022-07-21T19:29:13.078452Z","iopub.execute_input":"2022-07-21T19:29:13.079054Z","iopub.status.idle":"2022-07-21T19:29:13.110832Z","shell.execute_reply.started":"2022-07-21T19:29:13.079015Z","shell.execute_reply":"2022-07-21T19:29:13.109850Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# SAME WORK FOR test_data ","metadata":{}},{"cell_type":"code","source":"# Numerical :\n\ntest_data['LotFrontage'].fillna(test_data['LotFrontage'].median(),inplace=True)\ntest_data['MasVnrArea'].fillna(test_data['MasVnrArea'].median(),inplace=True)\n\n\n# Log Transformation: Applied on High Variance columns\n\ntest_data['LotArea'] = np.log(test_data['LotArea'])\n\n# Dropped Skewed Columns\n\nskewed_cols = ['LowQualFinSF','3SsnPorch','PoolArea','MiscVal','GarageYrBlt']\ntest_data.drop(skewed_cols,axis=1,inplace=True)\n\ntest_data.drop(['Id','GarageCars'],axis=1,inplace=True)\n\n# Categorical : \n\n\ndrop_col = ['Alley','PoolQC','MiscFeature','Fence']\ntest_data.drop(drop_col,axis=1,inplace=True)\n\ntest_data['MasVnrType'].fillna(test_data['MasVnrType'].mode()[0],inplace=True)\ntest_data['BsmtQual'].fillna(test_data['BsmtQual'].mode()[0],inplace=True)\ntest_data['BsmtCond'].fillna(test_data['BsmtCond'].mode()[0],inplace=True)\ntest_data['BsmtExposure'].fillna(test_data['BsmtExposure'].mode()[0],inplace=True)\ntest_data['BsmtFinType1'].fillna(test_data['BsmtFinType1'].mode()[0],inplace=True)\ntest_data['BsmtFinType2'].fillna(test_data['BsmtFinType2'].mode()[0],inplace=True)\ntest_data['FireplaceQu'].fillna(test_data['FireplaceQu'].mode()[0],inplace=True)\ntest_data['GarageType'].fillna(test_data['GarageType'].mode()[0],inplace=True)\ntest_data['GarageFinish'].fillna(test_data['GarageFinish'].mode()[0],inplace=True)\ntest_data['GarageQual'].fillna(test_data['GarageQual'].mode()[0],inplace=True)\ntest_data['GarageCond'].fillna(test_data['GarageCond'].mode()[0],inplace=True)\ntest_data['Electrical'].fillna(test_data['Electrical'].mode()[0],inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T19:29:18.087497Z","iopub.execute_input":"2022-07-21T19:29:18.087870Z","iopub.status.idle":"2022-07-21T19:29:18.123531Z","shell.execute_reply.started":"2022-07-21T19:29:18.087833Z","shell.execute_reply":"2022-07-21T19:29:18.122527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"le = LabelEncoder()\nfor col in cat_columns:\n    test_data[col] = le.fit_transform(test_data[col])","metadata":{"execution":{"iopub.status.busy":"2022-07-21T19:29:21.664803Z","iopub.execute_input":"2022-07-21T19:29:21.665474Z","iopub.status.idle":"2022-07-21T19:29:21.703905Z","shell.execute_reply.started":"2022-07-21T19:29:21.665438Z","shell.execute_reply":"2022-07-21T19:29:21.702994Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Some features of test_data had contain Nan so we apply the fillna function to fill it by median. we applied median because these are numerical columns**","metadata":{}},{"cell_type":"code","source":"test_data['BsmtFinSF1'].fillna(test_data['BsmtFinSF1'].median(),inplace=True)\ntest_data['BsmtUnfSF'].fillna(test_data['BsmtUnfSF'].median(),inplace=True)\ntest_data['TotalBsmtSF'].fillna(test_data['TotalBsmtSF'].median(),inplace=True)\ntest_data['BsmtFullBath'].fillna(test_data['BsmtFullBath'].median(),inplace=True)\ntest_data['GarageArea'].fillna(test_data['GarageArea'].median(),inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T19:29:29.473635Z","iopub.execute_input":"2022-07-21T19:29:29.473998Z","iopub.status.idle":"2022-07-21T19:29:29.486610Z","shell.execute_reply.started":"2022-07-21T19:29:29.473967Z","shell.execute_reply":"2022-07-21T19:29:29.485284Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data","metadata":{"execution":{"iopub.status.busy":"2022-07-21T08:57:35.340758Z","iopub.execute_input":"2022-07-21T08:57:35.341394Z","iopub.status.idle":"2022-07-21T08:57:35.380373Z","shell.execute_reply.started":"2022-07-21T08:57:35.341335Z","shell.execute_reply":"2022-07-21T08:57:35.379385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"input = train_data.drop(['SalePrice'],axis=1)\ntarget = train_data.SalePrice","metadata":{"execution":{"iopub.status.busy":"2022-07-21T19:29:34.660633Z","iopub.execute_input":"2022-07-21T19:29:34.661235Z","iopub.status.idle":"2022-07-21T19:29:34.669892Z","shell.execute_reply.started":"2022-07-21T19:29:34.661194Z","shell.execute_reply":"2022-07-21T19:29:34.668465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# BreakDown it into train and test part so we can train our Model or evaluate it","metadata":{}},{"cell_type":"markdown","source":"**we applied 80 and 20 percent which is applied by me after multiple trials**","metadata":{}},{"cell_type":"code","source":"x_train,x_test,y_train,y_test = train_test_split(input,target,test_size=0.2)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T19:29:37.767972Z","iopub.execute_input":"2022-07-21T19:29:37.770132Z","iopub.status.idle":"2022-07-21T19:29:37.778082Z","shell.execute_reply.started":"2022-07-21T19:29:37.770080Z","shell.execute_reply":"2022-07-21T19:29:37.777045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**I used XGBRegressor Model which is powerful one and applied done hyperparameter tunning **","metadata":{}},{"cell_type":"code","source":"para = {\n    'n_estimators': [300,200,100,400,500,600,700,900],\n    'max_depth' :[6,3,7,8,9,5,4],\n    'learning_rate' : [0.1,0.2,0.01,0.001,0.0001,0.02,0.002,0.5],\n    'subsample': [0.1,0.2,0.5,0.9,0.8,0.05],\n    'alpha': [10,14,20,22,30],\n    'booster' : ['gbtree','gblinear'],\n    'min_child_weight': [2,3,4,5,6,7,8,9],\n    'col_sample_bytree' : [0.5,0.6,0.55,0.85,0.68,0.9,1,0.7]\n}\n\nxg = xgb.XGBRegressor()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T19:29:40.012659Z","iopub.execute_input":"2022-07-21T19:29:40.013029Z","iopub.status.idle":"2022-07-21T19:29:40.023943Z","shell.execute_reply.started":"2022-07-21T19:29:40.012996Z","shell.execute_reply":"2022-07-21T19:29:40.019370Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**RandomizedSearchCV Model is applied to choose the best hyperparameters for XGBRegressor **","metadata":{}},{"cell_type":"code","source":"mod = RandomizedSearchCV(estimator=xg,param_distributions=para,cv = 5 , n_iter = 20)\nmod.fit(x_train,y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T19:29:42.763939Z","iopub.execute_input":"2022-07-21T19:29:42.764425Z","iopub.status.idle":"2022-07-21T19:30:45.000644Z","shell.execute_reply.started":"2022-07-21T19:29:42.764391Z","shell.execute_reply":"2022-07-21T19:30:44.999771Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mod.score(x_test,y_test)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T19:31:24.981803Z","iopub.execute_input":"2022-07-21T19:31:24.982183Z","iopub.status.idle":"2022-07-21T19:31:24.998771Z","shell.execute_reply.started":"2022-07-21T19:31:24.982153Z","shell.execute_reply":"2022-07-21T19:31:24.997309Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**These are the best Parameters for XGBRegressor on our data and according to our given parameters range**","metadata":{}},{"cell_type":"code","source":"mod.best_params_","metadata":{"execution":{"iopub.status.busy":"2022-07-21T19:31:28.268985Z","iopub.execute_input":"2022-07-21T19:31:28.269525Z","iopub.status.idle":"2022-07-21T19:31:28.278181Z","shell.execute_reply.started":"2022-07-21T19:31:28.269484Z","shell.execute_reply":"2022-07-21T19:31:28.276333Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred = mod.predict(test_data)\nsub = pd.read_csv('../input/home-data-for-ml-course/sample_submission.csv')\nsub['SalePrice'] = pred\nsub.to_csv('submission.csv',index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T19:31:31.832799Z","iopub.execute_input":"2022-07-21T19:31:31.833480Z","iopub.status.idle":"2022-07-21T19:31:31.868177Z","shell.execute_reply.started":"2022-07-21T19:31:31.833442Z","shell.execute_reply":"2022-07-21T19:31:31.867456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}