{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# DATA DESCRIPTION\n\n## Here's a brief version of what you'll find in the data description file.\n\n* SalePrice - the property's sale price in dollars. This is the target variable that you're trying to predict.\n* MSSubClass: The building class\n* MSZoning: The general zoning classification\n* LotFrontage: Linear feet of street connected to property\n* LotArea: Lot size in square feet\n* Street: Type of road access\n* Alley: Type of alley access\n* LotShape: General shape of property\n* LandContour: Flatness of the property\n* Utilities: Type of utilities available\n* LotConfig: Lot configuration\n* LandSlope: Slope of property\n* Neighborhood: Physical locations within Ames city limits\n* Condition1: Proximity to main road or railroad\n* Condition2: Proximity to main road or railroad (if a second is present)\n* BldgType: Type of dwelling\n* HouseStyle: Style of dwelling\n* OverallQual: Overall material and finish quality\n* OverallCond: Overall condition rating\n* YearBuilt: Original construction date\n* YearRemodAdd: Remodel date\n* RoofStyle: Type of roof\n* RoofMatl: Roof material\n* Exterior1st: Exterior covering on house\n* Exterior2nd: Exterior covering on house (if more than one material)\n* MasVnrType: Masonry veneer type\n* MasVnrArea: Masonry veneer area in square feet\n* ExterQual: Exterior material quality\n* ExterCond: Present condition of the material on the exterior\n* Foundation: Type of foundation\n* BsmtQual: Height of the basement\n* BsmtCond: General condition of the basement\n* BsmtExposure: Walkout or garden level basement walls\n* BsmtFinType1: Quality of basement finished area\n* BsmtFinSF1: Type 1 finished square feet\n* BsmtFinType2: Quality of second finished area (if present)\n* BsmtFinSF2: Type 2 finished square feet\n* BsmtUnfSF: Unfinished square feet of basement area\n* TotalBsmtSF: Total square feet of basement area\n* Heating: Type of heating\n* HeatingQC: Heating quality and condition\n* CentralAir: Central air conditioning\n* Electrical: Electrical system\n* 1stFlrSF: First Floor square feet\n* 2ndFlrSF: Second floor square feet\n* LowQualFinSF: Low quality finished square feet (all floors)\n* GrLivArea: Above grade (ground) living area square feet\n* BsmtFullBath: Basement full bathrooms\n* BsmtHalfBath: Basement half bathrooms\n* FullBath: Full bathrooms above grade\n* HalfBath: Half baths above grade\n* Bedroom: Number of bedrooms above basement level\n* Kitchen: Number of kitchens\n* KitchenQual: Kitchen quality\n* TotRmsAbvGrd: Total rooms above grade (does not include bathrooms)\n* Functional: Home functionality rating\n* Fireplaces: Number of fireplaces\n* FireplaceQu: Fireplace quality\n* GarageType: Garage location\n* GarageYrBlt: Year garage was built\n* GarageFinish: Interior finish of the garage\n* GarageCars: Size of garage in car capacity\n* GarageArea: Size of garage in square feet\n* GarageQual: Garage quality\n* GarageCond: Garage condition\n* PavedDrive: Paved driveway\n* WoodDeckSF: Wood deck area in square feet\n* OpenPorchSF: Open porch area in square feet\n* EnclosedPorch: Enclosed porch area in square feet\n* 3SsnPorch: Three season porch area in square feet\n* ScreenPorch: Screen porch area in square feet\n* PoolArea: Pool area in square feet\n* PoolQC: Pool quality\n* Fence: Fence quality\n* MiscFeature: Miscellaneous feature not covered in other categories\n* MiscVal: $Value of miscellaneous feature\n* MoSold: Month Sold\n* YrSold: Year Sold\n* SaleType: Type of sale\n* SaleCondition: Condition of sale","metadata":{}},{"cell_type":"markdown","source":"# Importing libraries","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport warnings\nwarnings.filterwarnings('ignore')\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn import metrics\nfrom sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score\nfrom sklearn.svm import SVR\nfrom sklearn.tree import DecisionTreeRegressor\nfrom sklearn.ensemble import RandomForestRegressor\nfrom sklearn import neighbors\nfrom math import sqrt\n%matplotlib inline","metadata":{"execution":{"iopub.status.busy":"2022-07-10T14:22:38.665108Z","iopub.execute_input":"2022-07-10T14:22:38.665782Z","iopub.status.idle":"2022-07-10T14:22:39.640267Z","shell.execute_reply.started":"2022-07-10T14:22:38.665645Z","shell.execute_reply":"2022-07-10T14:22:39.639256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Loading Train and Test datasets","metadata":{}},{"cell_type":"code","source":"df1= pd.read_csv('../input/house-prices-advanced-regression-techniques/train.csv')\ndf1","metadata":{"execution":{"iopub.status.busy":"2022-07-10T14:22:42.768938Z","iopub.execute_input":"2022-07-10T14:22:42.769839Z","iopub.status.idle":"2022-07-10T14:22:42.843861Z","shell.execute_reply.started":"2022-07-10T14:22:42.769799Z","shell.execute_reply":"2022-07-10T14:22:42.842734Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"valid = pd.read_csv('../input/house-prices-advanced-regression-techniques/test.csv')\nvalid","metadata":{"execution":{"iopub.status.busy":"2022-07-10T14:22:48.055292Z","iopub.execute_input":"2022-07-10T14:22:48.055677Z","iopub.status.idle":"2022-07-10T14:22:48.116447Z","shell.execute_reply.started":"2022-07-10T14:22:48.055646Z","shell.execute_reply":"2022-07-10T14:22:48.115328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Exploratory Data Analysis(EDA)","metadata":{}},{"cell_type":"code","source":"df1.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-10T14:22:54.723923Z","iopub.execute_input":"2022-07-10T14:22:54.725098Z","iopub.status.idle":"2022-07-10T14:22:54.750303Z","shell.execute_reply.started":"2022-07-10T14:22:54.725056Z","shell.execute_reply":"2022-07-10T14:22:54.749181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df1.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-10T14:22:57.314300Z","iopub.execute_input":"2022-07-10T14:22:57.314833Z","iopub.status.idle":"2022-07-10T14:22:57.414949Z","shell.execute_reply.started":"2022-07-10T14:22:57.314801Z","shell.execute_reply":"2022-07-10T14:22:57.413909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df1.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-10T14:23:00.588506Z","iopub.execute_input":"2022-07-10T14:23:00.588935Z","iopub.status.idle":"2022-07-10T14:23:00.615341Z","shell.execute_reply.started":"2022-07-10T14:23:00.588902Z","shell.execute_reply":"2022-07-10T14:23:00.614583Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df1.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-10T14:23:10.481691Z","iopub.execute_input":"2022-07-10T14:23:10.482138Z","iopub.status.idle":"2022-07-10T14:23:10.493929Z","shell.execute_reply.started":"2022-07-10T14:23:10.482100Z","shell.execute_reply":"2022-07-10T14:23:10.493268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df1.isnull().mean()","metadata":{"execution":{"iopub.status.busy":"2022-07-10T14:24:17.650767Z","iopub.execute_input":"2022-07-10T14:24:17.651746Z","iopub.status.idle":"2022-07-10T14:24:17.664618Z","shell.execute_reply.started":"2022-07-10T14:24:17.651707Z","shell.execute_reply":"2022-07-10T14:24:17.663527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Missing values percentage","metadata":{}},{"cell_type":"code","source":"def missing (df1):\n    # get the dictionary of column name and column values\n    missing_number = df1.isnull().sum().sort_values(ascending=False)\n    missing_percent = ( (df1.isnull().sum() / df1.isnull().count()) *100 ).sort_values(ascending=False)\n    missing_values = pd.concat([missing_number, missing_percent], axis=1, keys = ['Missing_Number', 'Missing_Percent'])\n#     missing_number = df1.isnull().sum().sort_values(ascending=False)\n#     missing_percent = ((df1.isnull().sum() / df1.isnull().count()) * 100).sort_values(ascending=False)\n#     missing_values = pd.concat([missing_number, missing_percent], axis=1, keys=['Missing_Number', 'Missing_Percent'])\n    return missing_values","metadata":{"execution":{"iopub.status.busy":"2022-07-10T14:37:59.522528Z","iopub.execute_input":"2022-07-10T14:37:59.523052Z","iopub.status.idle":"2022-07-10T14:37:59.530582Z","shell.execute_reply.started":"2022-07-10T14:37:59.523010Z","shell.execute_reply":"2022-07-10T14:37:59.529503Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"missing(df1)","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2022-07-10T14:38:01.303730Z","iopub.execute_input":"2022-07-10T14:38:01.304159Z","iopub.status.idle":"2022-07-10T14:38:01.334896Z","shell.execute_reply.started":"2022-07-10T14:38:01.304125Z","shell.execute_reply":"2022-07-10T14:38:01.334033Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Dropping the columns which have more than or equal to 40% of values as null ","metadata":{}},{"cell_type":"code","source":"for col in df1.columns:\n#     print(df1[col].isnull().mea//n())\n    if df1[col].isnull().mean()*100>40:\n        df1.drop(col,axis=1,inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-10T14:39:56.690301Z","iopub.execute_input":"2022-07-10T14:39:56.690960Z","iopub.status.idle":"2022-07-10T14:39:56.724936Z","shell.execute_reply.started":"2022-07-10T14:39:56.690914Z","shell.execute_reply":"2022-07-10T14:39:56.723916Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df1","metadata":{"execution":{"iopub.status.busy":"2022-07-10T14:40:21.013113Z","iopub.execute_input":"2022-07-10T14:40:21.013669Z","iopub.status.idle":"2022-07-10T14:40:21.041423Z","shell.execute_reply.started":"2022-07-10T14:40:21.013637Z","shell.execute_reply":"2022-07-10T14:40:21.040720Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df1.columns","metadata":{"execution":{"iopub.status.busy":"2022-07-10T14:40:22.842172Z","iopub.execute_input":"2022-07-10T14:40:22.842907Z","iopub.status.idle":"2022-07-10T14:40:22.850437Z","shell.execute_reply.started":"2022-07-10T14:40:22.842859Z","shell.execute_reply":"2022-07-10T14:40:22.849275Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df1.dtypes)\nprint('=========')\nprint(df1.dtypes.map(str))\nprint('======')\nprint(sns.countplot(df1.dtypes))\nprint(sns.countplot(df1.dtypes.map(str)))\nsns.countplot(df1.dtypes.map(str))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-10T14:46:26.322146Z","iopub.execute_input":"2022-07-10T14:46:26.322539Z","iopub.status.idle":"2022-07-10T14:46:26.508614Z","shell.execute_reply.started":"2022-07-10T14:46:26.322511Z","shell.execute_reply":"2022-07-10T14:46:26.507944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df1.dtypes.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-10T14:43:26.616324Z","iopub.execute_input":"2022-07-10T14:43:26.616795Z","iopub.status.idle":"2022-07-10T14:43:26.625486Z","shell.execute_reply.started":"2022-07-10T14:43:26.616759Z","shell.execute_reply":"2022-07-10T14:43:26.624340Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## year based filling of nullvalues in train dataset","metadata":{}},{"cell_type":"code","source":"# f = lambda x: x.median() if np.issubdtype(x.dtype, np.number) else x.mode().iloc[0]\nf = lambda x: x.median() if np.issubdtype(x.dtype, np.number) else x.mode().iloc[0]\nprint(f)\ndf1 = df1.fillna(df1.groupby('YrSold').transform(f))\n# df1 = df1.fillna(df1.groupby('YrSold').transform(f))\ndf1","metadata":{"execution":{"iopub.status.busy":"2022-07-10T15:14:16.349299Z","iopub.execute_input":"2022-07-10T15:14:16.349750Z","iopub.status.idle":"2022-07-10T15:14:16.513995Z","shell.execute_reply.started":"2022-07-10T15:14:16.349707Z","shell.execute_reply":"2022-07-10T15:14:16.512577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### finding q1,q2,q3,mean,median,mode,skewness,kurtosis","metadata":{}},{"cell_type":"code","source":"for col in df1.columns:\n    # process the int type of columns\n    if df1[col].dtypes != object:\n        q1 = df1[col].quantile(0.25)\n        q2 = df1[col].quantile(0.50)\n        q3 = df1[col].quantile(0.75)\n        IQR = q3 - q1\n        llp = q1-1.5*IQR\n        ulp = q3+1.5*IQR\n        print('column name',col)\n        print('q1',q1)\n        print('q2',q2)\n        print('q3',q3)\n        print('IQR',IQR)\n        print('llp',llp)\n        print('ulp',ulp)\n        print('mean:',df1[col].mean())\n        print('median:',df1[col].median())\n        print('mode',df1[col].mode()[0])\n        print('skewness:',df1[col].skew())\n        print('kurtosis:',df1[col].kurtosis())\n        print('std',df1[col].std())\n        print('max',df1[col].max())\n        print('min',df1[col].min())\n        print('null_value count:',df1[col].isnull().sum())\n        print('\\n')","metadata":{"execution":{"iopub.status.busy":"2022-07-10T15:16:16.272129Z","iopub.execute_input":"2022-07-10T15:16:16.272510Z","iopub.status.idle":"2022-07-10T15:16:16.427744Z","shell.execute_reply.started":"2022-07-10T15:16:16.272482Z","shell.execute_reply":"2022-07-10T15:16:16.426696Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df1.dtypes","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:10:39.03404Z","iopub.execute_input":"2022-06-24T15:10:39.034607Z","iopub.status.idle":"2022-06-24T15:10:39.043717Z","shell.execute_reply.started":"2022-06-24T15:10:39.034571Z","shell.execute_reply":"2022-06-24T15:10:39.042657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"raw","source":"for col in df1.columns:\n    if df1[col].dtypes == 'object':\n        print('column name:',col)\n        special = '[@_!#$%^&*()<>?/\\|}{~:-]'\n        print(df1[col].astype('str').str.count(special).sum())","metadata":{}},{"cell_type":"code","source":"df1['MSZoning'].unique()","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:10:39.045218Z","iopub.execute_input":"2022-06-24T15:10:39.045974Z","iopub.status.idle":"2022-06-24T15:10:39.057821Z","shell.execute_reply.started":"2022-06-24T15:10:39.045936Z","shell.execute_reply":"2022-06-24T15:10:39.056541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df1['RoofMatl'].unique()","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:10:39.059351Z","iopub.execute_input":"2022-06-24T15:10:39.060364Z","iopub.status.idle":"2022-06-24T15:10:39.068497Z","shell.execute_reply.started":"2022-06-24T15:10:39.060321Z","shell.execute_reply":"2022-06-24T15:10:39.067654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Q1 = df1.quantile(0.25)\nQ3 = df1.quantile(0.75)\nIQR = Q3 - Q1\nprint('outliers count of each columns')\n((df1 < (Q1 - 1.5 * IQR)) | (df1 > (Q3 + 1.5 * IQR))).sum()","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:10:39.069584Z","iopub.execute_input":"2022-06-24T15:10:39.069987Z","iopub.status.idle":"2022-06-24T15:10:39.135852Z","shell.execute_reply.started":"2022-06-24T15:10:39.06996Z","shell.execute_reply":"2022-06-24T15:10:39.134974Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data visualizations\n\n### Data visualisations using distplot,boxplot(because distplot and boxplot shows how data is distributed and if there are any outliers","metadata":{}},{"cell_type":"code","source":"count=1\nplt.subplots(figsize=(30,25))\nfor i in df1.columns:\n    if df1[i].dtypes!='object':\n        plt.subplot(6,7,count)\n        sns.distplot(df1[i])\n        count+=1\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:10:39.136974Z","iopub.execute_input":"2022-06-24T15:10:39.137274Z","iopub.status.idle":"2022-06-24T15:10:47.041396Z","shell.execute_reply.started":"2022-06-24T15:10:39.137246Z","shell.execute_reply":"2022-06-24T15:10:47.040432Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"count=1\nplt.subplots(figsize=(30,25))\nfor i in df1.columns:\n    if df1[i].dtypes!='object':\n        plt.subplot(6,7,count)\n        sns.boxplot(df1[i])\n        count+=1\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:10:47.042687Z","iopub.execute_input":"2022-06-24T15:10:47.043177Z","iopub.status.idle":"2022-06-24T15:10:50.496588Z","shell.execute_reply.started":"2022-06-24T15:10:47.043145Z","shell.execute_reply":"2022-06-24T15:10:50.495586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df1.dtypes","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:10:50.497994Z","iopub.execute_input":"2022-06-24T15:10:50.498343Z","iopub.status.idle":"2022-06-24T15:10:50.507452Z","shell.execute_reply.started":"2022-06-24T15:10:50.498297Z","shell.execute_reply":"2022-06-24T15:10:50.506436Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pip install autoviz","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:10:50.508829Z","iopub.execute_input":"2022-06-24T15:10:50.509196Z","iopub.status.idle":"2022-06-24T15:11:12.107824Z","shell.execute_reply.started":"2022-06-24T15:10:50.509163Z","shell.execute_reply":"2022-06-24T15:11:12.106454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from autoviz.AutoViz_Class import AutoViz_Class\nAV = AutoViz_Class()\ndf_av = AV.AutoViz('../input/house-prices-advanced-regression-techniques/train.csv')","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:11:12.109741Z","iopub.execute_input":"2022-06-24T15:11:12.110171Z","iopub.status.idle":"2022-06-24T15:12:21.501727Z","shell.execute_reply.started":"2022-06-24T15:11:12.110124Z","shell.execute_reply":"2022-06-24T15:12:21.500461Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Label encoding the train dataset","metadata":{}},{"cell_type":"code","source":"le=LabelEncoder()\nfor col in df1.columns:\n    if df1[col].dtypes == object:\n        df1[col]= le.fit_transform(df1[col])","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:12:21.504778Z","iopub.execute_input":"2022-06-24T15:12:21.506158Z","iopub.status.idle":"2022-06-24T15:12:23.469178Z","shell.execute_reply.started":"2022-06-24T15:12:21.506071Z","shell.execute_reply":"2022-06-24T15:12:23.468255Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Feature Selection","metadata":{}},{"cell_type":"code","source":"X=df1.drop('SalePrice',axis=1)\ny=df1['SalePrice']","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:12:23.475238Z","iopub.execute_input":"2022-06-24T15:12:23.475693Z","iopub.status.idle":"2022-06-24T15:12:23.485711Z","shell.execute_reply.started":"2022-06-24T15:12:23.475659Z","shell.execute_reply":"2022-06-24T15:12:23.484435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train test and split ","metadata":{}},{"cell_type":"code","source":"X_train,X_test,y_train,y_test = train_test_split(X,y,test_size=0.25,random_state=42)","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:12:23.487195Z","iopub.execute_input":"2022-06-24T15:12:23.488454Z","iopub.status.idle":"2022-06-24T15:12:23.498725Z","shell.execute_reply.started":"2022-06-24T15:12:23.488413Z","shell.execute_reply":"2022-06-24T15:12:23.496792Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Accuracies of different algorithms applied","metadata":{}},{"cell_type":"code","source":"def train_models(X_train, y_train):\n    \n #use Decision Tree\n   \n    tree = DecisionTreeRegressor(max_features=75,max_depth=4, random_state = 0)\n    tree.fit(X_train, y_train)\n    y_pred_tree = tree.predict(X_test)\n\n  #use the RandomForestRegressor\n    \n    rf = RandomForestRegressor(n_estimators = 100,max_features =75, random_state = 0)\n    rf.fit(X_train, y_train)\n    y_pred_rf= rf.predict(X_test)\n    \n  # use the support vector regressor\n    #from sklearn.svm import SVR\n    svr= SVR(kernel = 'rbf')\n    svr.fit(X_train, y_train)\n    y_pred_svr = svr.predict(X_test)\n    \n    #from sklearn.svm import SVR\n    svr_l= SVR(kernel = 'linear')\n    svr_l.fit(X_train, y_train)\n    y_pred_svr_linear = svr_l.predict(X_test)\n\n  # use the knn regressor\n    knn = neighbors.KNeighborsRegressor()\n    knn.fit(X_train, y_train)\n    y_pred_knn = knn.predict(X_test)\n    \n  # metrics of decision tree regressor\n    meanAbErr_tree= metrics.mean_absolute_error(y_test, y_pred_tree)\n    meanSqErr_tree= metrics.mean_squared_error(y_test, y_pred_tree)\n    rootMeanSqErr_tree= np.sqrt(metrics.mean_squared_error(y_test, y_pred_tree))\n\n  # metrics of random forest regressor\n    meanAbErr_rf= metrics.mean_absolute_error(y_test, y_pred_rf)\n    meanSqErr_rf= metrics.mean_squared_error(y_test, y_pred_rf)\n    rootMeanSqErr_rf= np.sqrt(metrics.mean_squared_error(y_test, y_pred_rf))\n  \n  # metrics of knn regressor\n    meanAbErr_knn = metrics.mean_absolute_error(y_test, y_pred_knn)\n    meanSqErr_knn = metrics.mean_squared_error(y_test, y_pred_knn)\n    rootMeanSqErr_knn= np.sqrt(metrics.mean_squared_error(y_test, y_pred_knn)) \n\n  # metrics of svr regressor\n    meanAbErr_svr = metrics.mean_absolute_error(y_test, y_pred_svr_linear)\n    meanSqErr_svr = metrics.mean_squared_error(y_test, y_pred_svr_linear)\n    rootMeanSqErr_svr= np.sqrt(metrics.mean_squared_error(y_test, y_pred_svr_linear)) \n\n  #print the tranning accurancy of each model\n\n    print('[1]Decision Tree Training Accurancy: ', r2_score(y_test,y_pred_tree))\n    print('Mean Absolute Error:', meanAbErr_tree)\n    print('Mean Square Error:', meanSqErr_tree)\n    print('Root Mean Square Error:', rootMeanSqErr_tree)\n    print('\\t')\n    print('[2]RandomForestRegressor Training Accurancy: ',r2_score(y_test,y_pred_rf))\n    print('Mean Absolute Error:', meanAbErr_rf)\n    print('Mean Square Error:', meanSqErr_rf)\n    print('Root Mean Square Error:', rootMeanSqErr_rf)\n    print('\\t')    \n    print('[3]SupportvectorRegression Accuracy(rbf): ', r2_score(y_test,y_pred_svr))\n    print('\\t')\n    print('[4]SupportvectorRegression Accuracy(linear): ', r2_score(y_test,y_pred_svr_linear))\n    print('Mean Absolute Error:', meanAbErr_svr)\n    print('Mean Square Error:', meanSqErr_svr)\n    print('Root Mean Square Error:', rootMeanSqErr_svr)\n    print('\\t')\n    print('[5]knn Training Accurancy: ', r2_score(y_test,y_pred_knn))\n    print('Mean Absolute Error:', meanAbErr_knn)\n    print('Mean Square Error:', meanSqErr_knn)\n    print('Root Mean Square Error:', rootMeanSqErr_knn)\n    print('\\t')\n    \n","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:12:23.500494Z","iopub.execute_input":"2022-06-24T15:12:23.501532Z","iopub.status.idle":"2022-06-24T15:12:23.521279Z","shell.execute_reply.started":"2022-06-24T15:12:23.501481Z","shell.execute_reply":"2022-06-24T15:12:23.520369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_models(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:12:23.522476Z","iopub.execute_input":"2022-06-24T15:12:23.523371Z","iopub.status.idle":"2022-06-24T15:12:37.542396Z","shell.execute_reply.started":"2022-06-24T15:12:23.523324Z","shell.execute_reply":"2022-06-24T15:12:37.538493Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### User defined function is for showing purpose only that's why\n### We can't use user defined function for model prediction for validation\n### so that why we have to apply algorithms seperately, iam using random forest because it gets good accuracy","metadata":{}},{"cell_type":"markdown","source":"# Multiple Linear Regression","metadata":{}},{"cell_type":"code","source":"from sklearn.linear_model import LinearRegression\nmlr = LinearRegression()  \nmlr.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:12:37.546577Z","iopub.execute_input":"2022-06-24T15:12:37.547432Z","iopub.status.idle":"2022-06-24T15:12:37.606508Z","shell.execute_reply.started":"2022-06-24T15:12:37.547378Z","shell.execute_reply":"2022-06-24T15:12:37.605224Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred_mlr= mlr.predict(X_test)\ny_pred_mlr","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:12:37.6091Z","iopub.execute_input":"2022-06-24T15:12:37.611101Z","iopub.status.idle":"2022-06-24T15:12:37.669982Z","shell.execute_reply.started":"2022-06-24T15:12:37.611047Z","shell.execute_reply":"2022-06-24T15:12:37.667863Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"r2_mlr =r2_score(y_test,y_pred_mlr)\nprint('r2_score:',r2_mlr*100)","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:12:37.673447Z","iopub.execute_input":"2022-06-24T15:12:37.676695Z","iopub.status.idle":"2022-06-24T15:12:37.694765Z","shell.execute_reply.started":"2022-06-24T15:12:37.676641Z","shell.execute_reply":"2022-06-24T15:12:37.692568Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# validation dataset\n\n### Same EDA and preprocessing steps have to be followed for validation dataset same as for train dataset","metadata":{}},{"cell_type":"code","source":"valid","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:12:37.698359Z","iopub.execute_input":"2022-06-24T15:12:37.701937Z","iopub.status.idle":"2022-06-24T15:12:37.7573Z","shell.execute_reply.started":"2022-06-24T15:12:37.701863Z","shell.execute_reply":"2022-06-24T15:12:37.756571Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"missing(valid)","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:12:37.758536Z","iopub.execute_input":"2022-06-24T15:12:37.758935Z","iopub.status.idle":"2022-06-24T15:12:37.799732Z","shell.execute_reply.started":"2022-06-24T15:12:37.758908Z","shell.execute_reply":"2022-06-24T15:12:37.798892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Dropping nullvalues of more than 40% for validation dataset","metadata":{}},{"cell_type":"code","source":"for col in valid.columns:\n    if valid[col].isnull().mean()*100>40:\n        valid.drop(col,axis=1,inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:12:37.801113Z","iopub.execute_input":"2022-06-24T15:12:37.801644Z","iopub.status.idle":"2022-06-24T15:12:37.847059Z","shell.execute_reply.started":"2022-06-24T15:12:37.801606Z","shell.execute_reply":"2022-06-24T15:12:37.846367Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"valid","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:12:37.848246Z","iopub.execute_input":"2022-06-24T15:12:37.848718Z","iopub.status.idle":"2022-06-24T15:12:37.879545Z","shell.execute_reply.started":"2022-06-24T15:12:37.848689Z","shell.execute_reply":"2022-06-24T15:12:37.878521Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"f = lambda x: x.median() if np.issubdtype(x.dtype, np.number) else x.mode().iloc[0]\nvalid = valid.fillna(valid.groupby('YrSold').transform(f))\nvalid","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:12:37.881291Z","iopub.execute_input":"2022-06-24T15:12:37.881833Z","iopub.status.idle":"2022-06-24T15:12:38.057742Z","shell.execute_reply.started":"2022-06-24T15:12:37.881785Z","shell.execute_reply":"2022-06-24T15:12:38.056838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"valid.columns","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:12:38.059038Z","iopub.execute_input":"2022-06-24T15:12:38.05983Z","iopub.status.idle":"2022-06-24T15:12:38.068142Z","shell.execute_reply.started":"2022-06-24T15:12:38.059795Z","shell.execute_reply":"2022-06-24T15:12:38.066925Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"le=LabelEncoder()\nfor col in valid.columns:\n    if valid[col].dtypes == 'object':\n        valid[col]= le.fit_transform(valid[col])","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:12:38.069795Z","iopub.execute_input":"2022-06-24T15:12:38.070318Z","iopub.status.idle":"2022-06-24T15:12:38.137761Z","shell.execute_reply.started":"2022-06-24T15:12:38.070259Z","shell.execute_reply":"2022-06-24T15:12:38.136598Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"valid['MSZoning'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:12:38.139715Z","iopub.execute_input":"2022-06-24T15:12:38.140243Z","iopub.status.idle":"2022-06-24T15:12:38.149378Z","shell.execute_reply.started":"2022-06-24T15:12:38.140195Z","shell.execute_reply":"2022-06-24T15:12:38.148016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"valid","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:12:38.151534Z","iopub.execute_input":"2022-06-24T15:12:38.152255Z","iopub.status.idle":"2022-06-24T15:12:38.185657Z","shell.execute_reply.started":"2022-06-24T15:12:38.152194Z","shell.execute_reply":"2022-06-24T15:12:38.184586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_valid = mlr.predict(valid)","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2022-06-24T15:12:38.18708Z","iopub.execute_input":"2022-06-24T15:12:38.187454Z","iopub.status.idle":"2022-06-24T15:12:38.204464Z","shell.execute_reply.started":"2022-06-24T15:12:38.187415Z","shell.execute_reply":"2022-06-24T15:12:38.203342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Validation data prediction","metadata":{}},{"cell_type":"code","source":"y_valid","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:12:38.209185Z","iopub.execute_input":"2022-06-24T15:12:38.210866Z","iopub.status.idle":"2022-06-24T15:12:38.22542Z","shell.execute_reply.started":"2022-06-24T15:12:38.210806Z","shell.execute_reply":"2022-06-24T15:12:38.224391Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"output = pd.DataFrame({\"Id\": valid['Id'],\"SalePrice\": y_valid})\noutput","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:12:38.226897Z","iopub.execute_input":"2022-06-24T15:12:38.232762Z","iopub.status.idle":"2022-06-24T15:12:38.259737Z","shell.execute_reply.started":"2022-06-24T15:12:38.2327Z","shell.execute_reply":"2022-06-24T15:12:38.258011Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Save the output\noutput.to_csv(\"submission5.csv\", index=False)\noutput.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-06-24T15:12:38.262359Z","iopub.execute_input":"2022-06-24T15:12:38.263223Z","iopub.status.idle":"2022-06-24T15:12:38.299332Z","shell.execute_reply.started":"2022-06-24T15:12:38.263162Z","shell.execute_reply":"2022-06-24T15:12:38.298228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}