{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-08T02:35:02.358480Z","iopub.execute_input":"2022-07-08T02:35:02.359345Z","iopub.status.idle":"2022-07-08T02:35:02.390291Z","shell.execute_reply.started":"2022-07-08T02:35:02.359243Z","shell.execute_reply":"2022-07-08T02:35:02.389355Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"house=pd.read_csv('../input/house-prices-advanced-regression-techniques/train.csv')\nhouse.head()\n                  ","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:02.391730Z","iopub.execute_input":"2022-07-08T02:35:02.392234Z","iopub.status.idle":"2022-07-08T02:35:02.461334Z","shell.execute_reply.started":"2022-07-08T02:35:02.392204Z","shell.execute_reply":"2022-07-08T02:35:02.460151Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"house.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:02.463724Z","iopub.execute_input":"2022-07-08T02:35:02.464104Z","iopub.status.idle":"2022-07-08T02:35:02.497189Z","shell.execute_reply.started":"2022-07-08T02:35:02.464055Z","shell.execute_reply":"2022-07-08T02:35:02.496262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp=pd.DataFrame(house.isna().sum(),columns=(['number']))\ntemp2=temp[temp['number']>0]                \ntemp2.sort_values(['number'],ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:02.498133Z","iopub.execute_input":"2022-07-08T02:35:02.498833Z","iopub.status.idle":"2022-07-08T02:35:02.517126Z","shell.execute_reply.started":"2022-07-08T02:35:02.498797Z","shell.execute_reply":"2022-07-08T02:35:02.516113Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So we have a bunch of variables here. Let;s first narrow down on those thta have a big number of Na values. \n1. We have a big number of na valrus for Pool QC. Basedo nthe data dictionary provided, Na refers to No pool ion the house. We should change the value accordingly later. \n2. MiscFeature --> na again refers to NOne or no misc feature \n3. Alley --> na refers to no alley access \n4. Fence --> na refers to fence \n5. FireplaceQU --> na rfers to non fireplace \n6. Lot frontage --> this is a cotinuos number so we might have to some data imputation later on \n7. the garage relatged ones --> we realise they all have the same numebr of na values. So it is clean in that sense, just that na refers to no garage in the house \n8. For the basement related ones, some have 38 missing values while some have 37. not sure why there is a discrpency here. BUt we will do our clean up and investgauite later if there is a need to. \n9. The masonry veneer related ones, there was no na descrption in the dictiopnary. So we can either assume there is no masonry veneer or we could remove them. Let's assume that they dont have the masonry thing and impute the value accoridngly. \n10. Finally for electrical, seems to be a missing completely at random issue, but we can investigate later. \n\nNow let's do some clean up for majority of the columns and see again whatr we are working with. \n\n\nFor the values that have a large number of NA values, (>80%), it might nto make sense to keep them since the variabels might not be useful. We will also remove the accopmanying variables if any as well since this will likley skew the data as well.  ","metadata":{}},{"cell_type":"code","source":"house=house.drop(['PoolArea','PoolQC','MiscFeature','MiscVal','Alley','Fence',\n                  'FireplaceQu','Fireplaces','LotFrontage','Id'],axis=1)\n ","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:02.518881Z","iopub.execute_input":"2022-07-08T02:35:02.519223Z","iopub.status.idle":"2022-07-08T02:35:02.526229Z","shell.execute_reply.started":"2022-07-08T02:35:02.519192Z","shell.execute_reply":"2022-07-08T02:35:02.525006Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"When we focus on the garage related fields, one of which is the year that the garaage is built. since the values will be missing, we shall take out the whole column. We will thyem imput the rest accoridngly.  ","metadata":{}},{"cell_type":"code","source":"house=house.drop(['GarageYrBlt'],axis=1)\nhouse['GarageType']=house['GarageType'].fillna('no_g')\nhouse['GarageFinish']=house['GarageFinish'].fillna('no_g')\nhouse['GarageQual']=house['GarageQual'].fillna('no_g')\nhouse['GarageCond']=house['GarageCond'].fillna('no_g')\nhouse['GarageType']=house['GarageType'].fillna('no_g')\nhouse['BsmtExposure']=house['BsmtExposure'].fillna('no_b')    \nhouse['BsmtFinType1']=house['BsmtFinType1'].fillna('no_b')    \nhouse['BsmtFinType2']=house['BsmtFinType2'].fillna('no_b')    \nhouse['BsmtCond']=house['BsmtCond'].fillna('no_b')    \nhouse['BsmtQual']=house['BsmtQual'].fillna('no_b')    \nhouse['MasVnrType']=house['MasVnrType'].fillna('no_m')\nhouse['MasVnrArea']=house['MasVnrArea'].fillna(0)\n\n\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:02.527415Z","iopub.execute_input":"2022-07-08T02:35:02.527790Z","iopub.status.idle":"2022-07-08T02:35:02.553006Z","shell.execute_reply.started":"2022-07-08T02:35:02.527759Z","shell.execute_reply":"2022-07-08T02:35:02.551983Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp=pd.DataFrame(house.isna().sum(),columns=(['number']))\ntemp2=temp[temp['number']>0]                \ntemp2.sort_values(['number'],ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:02.554909Z","iopub.execute_input":"2022-07-08T02:35:02.555833Z","iopub.status.idle":"2022-07-08T02:35:02.573334Z","shell.execute_reply.started":"2022-07-08T02:35:02.555798Z","shell.execute_reply":"2022-07-08T02:35:02.572252Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let;s check out this Electrical one. ","metadata":{}},{"cell_type":"code","source":"test=house[house['Electrical'].isna()]\ntest[['Electrical','Utilities']]","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:02.575538Z","iopub.execute_input":"2022-07-08T02:35:02.575976Z","iopub.status.idle":"2022-07-08T02:35:02.588739Z","shell.execute_reply.started":"2022-07-08T02:35:02.575936Z","shell.execute_reply":"2022-07-08T02:35:02.587221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"house['Electrical'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:02.595001Z","iopub.execute_input":"2022-07-08T02:35:02.597367Z","iopub.status.idle":"2022-07-08T02:35:02.607261Z","shell.execute_reply.started":"2022-07-08T02:35:02.597325Z","shell.execute_reply":"2022-07-08T02:35:02.605925Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We shall impute it with the mode since the Utilities mentioned that that the house hs access to all pub.","metadata":{}},{"cell_type":"code","source":"house['Electrical']=house['Electrical'].fillna('SBrkr')","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:02.611711Z","iopub.execute_input":"2022-07-08T02:35:02.612214Z","iopub.status.idle":"2022-07-08T02:35:02.617177Z","shell.execute_reply.started":"2022-07-08T02:35:02.612182Z","shell.execute_reply":"2022-07-08T02:35:02.616298Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now that we have dealt with the NA values and removed the columns for the first round, leet;s convert our cat varriables to cont variables first.\n\nWee will do this via nominal encoding as well as frequieency ecoding, since we dont have alot of differnt options per column. \n\nBut Before that, let;s split our dataset with a train test split, so that we dopnt run the risk of overfitting and prevent any data leakage. \n\n[edit] After which let;s scale our cotnous vairables frist using standard scaler. ","metadata":{}},{"cell_type":"code","source":"house['SalePrice']","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:02.625758Z","iopub.execute_input":"2022-07-08T02:35:02.626696Z","iopub.status.idle":"2022-07-08T02:35:02.636234Z","shell.execute_reply.started":"2022-07-08T02:35:02.626646Z","shell.execute_reply":"2022-07-08T02:35:02.635236Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"house.isna()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:02.637826Z","iopub.execute_input":"2022-07-08T02:35:02.638533Z","iopub.status.idle":"2022-07-08T02:35:02.673694Z","shell.execute_reply.started":"2022-07-08T02:35:02.638491Z","shell.execute_reply":"2022-07-08T02:35:02.672580Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split \n\ntarget=house['SalePrice']\nfeatures=house.drop('SalePrice',axis=1)\ntrain_x,test_x,train_y,test_y=train_test_split(house.drop('SalePrice',axis=1),house['SalePrice'],test_size=0.2)\n\n\n\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:02.675635Z","iopub.execute_input":"2022-07-08T02:35:02.675976Z","iopub.status.idle":"2022-07-08T02:35:03.987582Z","shell.execute_reply.started":"2022-07-08T02:35:02.675945Z","shell.execute_reply":"2022-07-08T02:35:03.986265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.preprocessing import OrdinalEncoder\n\nnumeric_data = train_x.select_dtypes(include=[np.number])\ncategorical_data = train_x.select_dtypes(exclude=[np.number])\n\n\ndef encode(dataset):\n    temp=[]\n    frequency=[]\n    for x in dataset.columns:\n    \n        if x.endswith('Qual'):\n            temp.append(x)\n        elif x.endswith('Cond'):\n            temp.append(x)\n\n        elif x.endswith('QC'):\n            temp.append(x)\n\n        else: \n            frequency.append(x)\n\n        \n    encoder=OrdinalEncoder()\n    encoder.fit(dataset[temp])\n    dataset[temp]=encoder.transform(dataset[temp])\n    \n    for y in frequency:\n        temp=dict(dataset[y].value_counts())\n        dataset[y]=dataset[y].replace(temp)\n    return dataset\n\n\ntrain_x_cat=encode(categorical_data)\n\ntrain_x_cat = train_x_cat.reset_index()\nnumeric_data=numeric_data.reset_index()\n\n\ntrain_x=pd.concat([numeric_data,train_x_cat],axis=1)\ntrain_x=train_x.drop('index',axis=1)\ntrain_x.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:03.990983Z","iopub.execute_input":"2022-07-08T02:35:03.991483Z","iopub.status.idle":"2022-07-08T02:35:04.199195Z","shell.execute_reply.started":"2022-07-08T02:35:03.991407Z","shell.execute_reply":"2022-07-08T02:35:04.198130Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nnumeric_data = test_x.select_dtypes(include=[np.number])\ncategorical_data = test_x.select_dtypes(exclude=[np.number])\n\n\ndef encode(dataset):\n    temp=[]\n    frequency=[]\n    for x in dataset.columns:\n    \n        if x.endswith('Qual'):\n            temp.append(x)\n        elif x.endswith('Cond'):\n            temp.append(x)\n\n        elif x.endswith('QC'):\n            temp.append(x)\n\n        else: \n            frequency.append(x)\n\n        \n    encoder=OrdinalEncoder()\n    encoder.fit(dataset[temp])\n    dataset[temp]=encoder.transform(dataset[temp])\n    \n    for y in frequency:\n        temp=dict(dataset[y].value_counts())\n        dataset[y]=dataset[y].replace(temp)\n    return dataset\n\n\ntrain_x_cat=encode(categorical_data)\n\ntrain_x_cat = train_x_cat.reset_index()\nnumeric_data=numeric_data.reset_index()\n\n\ntest_x=pd.concat([numeric_data,train_x_cat],axis=1)\ntest_x=test_x.drop('index',axis=1)\ntest_x.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:04.201914Z","iopub.execute_input":"2022-07-08T02:35:04.202605Z","iopub.status.idle":"2022-07-08T02:35:04.338035Z","shell.execute_reply.started":"2022-07-08T02:35:04.202557Z","shell.execute_reply":"2022-07-08T02:35:04.337012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Nicee. Now we can do further cleaning. We shall first start with variacne threshold to removee columns that have many similar values whihc will nto be useful for our predictions. ","metadata":{}},{"cell_type":"code","source":"\n\n\nfrom sklearn.feature_selection import VarianceThreshold\n#def fucntion for the filtering and retianign of column names \ndef variance_threshold_selector(data,threshold=0.5):\n    selector = VarianceThreshold(threshold)\n    selector.fit(data)\n    return data[data.columns[selector.get_support(indices=True)]]\n\n\n\ntrain_x=variance_threshold_selector(train_x)\ntest_x=test_x[list(train_x.columns)]\ntrain_x.info()\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:04.339645Z","iopub.execute_input":"2022-07-08T02:35:04.340351Z","iopub.status.idle":"2022-07-08T02:35:04.538972Z","shell.execute_reply.started":"2022-07-08T02:35:04.340305Z","shell.execute_reply":"2022-07-08T02:35:04.537609Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We manmaged to gety rid of 10 addtional columns. Now we shall do a correlation compairson and remove above say.. an abs corr of 0.8. \n\nAfter this, we can do another correlation test with respect to the target salesprice to see if we can observe more important varuiables that could be important","metadata":{}},{"cell_type":"code","source":"house_corr=train_x.corr()\n\n#create a list to append columns that have a high correlation \nlist_remove=[]\n\nfor x in house_corr.columns:\n    for y in house_corr.columns:\n        if np.abs(house_corr[x][y])>0.7:\n            #not to include itself since corr will be 1 \n            if house_corr[x][y]!=1:\n                list_remove.append(y)\n    \n\n\n#remove duplciates from list_remove\n# to remove duplicated from list \nresult = [] \n[result.append(x) for x in list_remove if x not in result] \n\n\nresult\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:04.540650Z","iopub.execute_input":"2022-07-08T02:35:04.541292Z","iopub.status.idle":"2022-07-08T02:35:04.621290Z","shell.execute_reply.started":"2022-07-08T02:35:04.541251Z","shell.execute_reply":"2022-07-08T02:35:04.620471Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_x=train_x.drop(result,axis=1)\ntest_x=test_x[list(train_x.columns)]","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:04.622414Z","iopub.execute_input":"2022-07-08T02:35:04.622900Z","iopub.status.idle":"2022-07-08T02:35:04.629187Z","shell.execute_reply.started":"2022-07-08T02:35:04.622870Z","shell.execute_reply":"2022-07-08T02:35:04.628339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp_corr=pd.concat([train_x,train_y],axis=1)\nhouse_corr_2=temp_corr.corr()\ntemp=house_corr_2['SalePrice'].sort_values(ascending=False)[1:21]\ncorr_list=pd.DataFrame(temp)\ncorr_lists=list(corr_list.index)\n#corr_lists=list(corr_list['test'])\nprint(corr_list)\n\n\n#create a list to append columns that have a high correlation \n#list_remove=[]\n\n#for x in house_corr.columns:\n#    for y in house_corr.columns:\n#        if np.abs(house_corr[x][y])>0.7:\n#            #not to include itself since corr will be 1 \n#            if house_corr[x][y]!=1:\n#                list_remove.append(y)\n    \n\n\n#remove duplciates from list_remove\n# to remove duplicated from list \n#result = [] \n#[result.append(x) for x in list_remove if x not in result] \n\n\n#result\n","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:04.630430Z","iopub.execute_input":"2022-07-08T02:35:04.631003Z","iopub.status.idle":"2022-07-08T02:35:04.654346Z","shell.execute_reply.started":"2022-07-08T02:35:04.630968Z","shell.execute_reply":"2022-07-08T02:35:04.653207Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's keep this in mind first. The thing i want to try out here for now is to identify using difference feautre seelction methods, whether we get similar top features to select from. \n\nNext up, let's try using the chi sqaure method ","metadata":{}},{"cell_type":"code","source":"from sklearn.feature_selection import SelectKBest, chi2,mutual_info_classif\n\nchi2=SelectKBest(chi2,k=12).fit(train_x,train_y)\n\n\ndfscores=pd.DataFrame(chi2.scores_,columns=[\"Score\"])\ndfcolumns=pd.DataFrame(train_x.columns)\nfeatures_rank=pd.concat([dfcolumns,dfscores],axis=1)\nfeatures_rank.columns=['Features','Score']\nchi2_features=list(features_rank['Features'][0:20])\nprint(chi2_features)\nfeatures_rank.nlargest(12,columns='Score')\n \n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:04.657297Z","iopub.execute_input":"2022-07-08T02:35:04.657673Z","iopub.status.idle":"2022-07-08T02:35:04.723591Z","shell.execute_reply.started":"2022-07-08T02:35:04.657637Z","shell.execute_reply":"2022-07-08T02:35:04.722126Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now let;s try using mutual info gain ","metadata":{}},{"cell_type":"code","source":"mutual_info=mutual_info_classif(train_x,train_y)\nmutual_data=pd.Series(mutual_info,index=train_x.columns)\nmutual_data.sort_values(ascending=False)[0:10]\nmutual_data_train=list(mutual_data.index[0:20])\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:04.725340Z","iopub.execute_input":"2022-07-08T02:35:04.726544Z","iopub.status.idle":"2022-07-08T02:35:10.340369Z","shell.execute_reply.started":"2022-07-08T02:35:04.726490Z","shell.execute_reply":"2022-07-08T02:35:10.339261Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lastly, let's try Feature Importance ","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import ExtraTreesClassifier\nimport matplotlib.pyplot as plt\nmodel=ExtraTreesClassifier()\nmodel.fit(train_x,train_y)\n\n\nranked_features=pd.Series(model.feature_importances_,index=train_x.columns)\nranked_features.nlargest(12).plot(kind='barh')\nfeat_imp=list(ranked_features.index[0:20])\nplt.show()\n\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:10.341703Z","iopub.execute_input":"2022-07-08T02:35:10.342123Z","iopub.status.idle":"2022-07-08T02:35:13.391975Z","shell.execute_reply.started":"2022-07-08T02:35:10.342082Z","shell.execute_reply":"2022-07-08T02:35:13.390898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Pre processing step after looking at the various tests ","metadata":{}},{"cell_type":"markdown","source":"And for test set ","metadata":{}},{"cell_type":"markdown","source":"SO now that we have tried all three types of methods and gotten some outputs for each type of variable seelction tehcnique, what we will do is to segregtae three differnt trainign sets basedo n theb est features ot select basedo n our different varibale seleciton methods and see what measures we get. \n\nWe will try the following algorithms: \n1. Linear Regression \n2. lasso (since this 'automatically' performs variabel selection, we can use the full dataset for this protion. \n3. Ridge Regression ","metadata":{}},{"cell_type":"code","source":"#prep the sets first \n\n# for the corr \ncorr_train_x=train_x[corr_lists]\ncorr_test_x=test_x[corr_lists]\n\n#for the chi2\nchi2_train_x=train_x[chi2_features]\nchi2_test_x=test_x[chi2_features]\n\n#for the mutual info \nmut_train_x=train_x[mutual_data_train]\nmut_test_x=test_x[mutual_data_train]\n\n# for the feture importance \nfeatimp_train_x=train_x[feat_imp]\nfeatimp_test_x=test_x[feat_imp]\n\n\n\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:13.393332Z","iopub.execute_input":"2022-07-08T02:35:13.393677Z","iopub.status.idle":"2022-07-08T02:35:13.404265Z","shell.execute_reply.started":"2022-07-08T02:35:13.393645Z","shell.execute_reply":"2022-07-08T02:35:13.403499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.linear_model import LinearRegression,RidgeCV,LassoCV,ElasticNet\nfrom sklearn.ensemble import RandomForestRegressor\nfrom sklearn import metrics\n","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:13.405402Z","iopub.execute_input":"2022-07-08T02:35:13.405748Z","iopub.status.idle":"2022-07-08T02:35:13.419525Z","shell.execute_reply.started":"2022-07-08T02:35:13.405717Z","shell.execute_reply":"2022-07-08T02:35:13.418248Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def model(model,trainx,trainy,testx,testy):\n    instance=model\n    reg=instance.fit(trainx,trainy)\n    y_pred=reg.predict(testx)\n    print('r-sqaured score is:'+ str(metrics.r2_score(testy,y_pred)))\n    print('MSE:'+ str(metrics.mean_squared_error(testy,y_pred)))\n    print('MAE:'+ str(metrics.mean_absolute_error(testy,y_pred)))\n    ","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:13.420917Z","iopub.execute_input":"2022-07-08T02:35:13.421670Z","iopub.status.idle":"2022-07-08T02:35:13.477117Z","shell.execute_reply.started":"2022-07-08T02:35:13.421623Z","shell.execute_reply":"2022-07-08T02:35:13.476176Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 1. Normal Linear Regression ","metadata":{}},{"cell_type":"code","source":"#raw dataset as well \nmodel(LinearRegression(),train_x,train_y,test_x,test_y)\n\n#corr\nmodel(LinearRegression(),corr_train_x,train_y,corr_test_x,test_y)\n\n#chi2\nmodel(LinearRegression(),chi2_train_x,train_y,chi2_test_x,test_y)\n\n#mutual info \nmodel(LinearRegression(),mut_train_x,train_y,mut_test_x,test_y)\n\n#feature importance\nmodel(LinearRegression(),featimp_train_x,train_y,featimp_test_x,test_y)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:13.478257Z","iopub.execute_input":"2022-07-08T02:35:13.479187Z","iopub.status.idle":"2022-07-08T02:35:13.592009Z","shell.execute_reply.started":"2022-07-08T02:35:13.479152Z","shell.execute_reply":"2022-07-08T02:35:13.590705Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 2. Ridge ","metadata":{}},{"cell_type":"code","source":"#raw dataset as well \nmodel(RidgeCV(alphas=[1e-3, 1e-2, 1e-1, 1]),train_x,train_y,test_x,test_y)\n\n#corr\nmodel(RidgeCV(alphas=[1e-3, 1e-2, 1e-1, 1]),corr_train_x,train_y,corr_test_x,test_y)\n\n#chi2\nmodel(RidgeCV(alphas=[1e-3, 1e-2, 1e-1, 1]),chi2_train_x,train_y,chi2_test_x,test_y)\n\n#mutual info \nmodel(RidgeCV(alphas=[1e-3, 1e-2, 1e-1, 1]),mut_train_x,train_y,mut_test_x,test_y)\n\n#feature importance\nmodel(RidgeCV(alphas=[1e-3, 1e-2, 1e-1, 1]),featimp_train_x,train_y,featimp_test_x,test_y)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:13.593959Z","iopub.execute_input":"2022-07-08T02:35:13.595827Z","iopub.status.idle":"2022-07-08T02:35:13.721569Z","shell.execute_reply.started":"2022-07-08T02:35:13.595777Z","shell.execute_reply":"2022-07-08T02:35:13.720374Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 3. Lasso ","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import StratifiedKFold\n\ncv=StratifiedKFold(n_splits=5, random_state=None, shuffle=False)\n\n\n#raw dataset as well \nmodel(LassoCV(cv=cv, random_state=0,alphas=[1e-3, 1e-2, 1e-1, 1]),train_x,train_y,test_x,test_y)\n\n#corr\nmodel(LassoCV(cv=cv, random_state=0,alphas=[1e-3, 1e-2, 1e-1, 1]),corr_train_x,train_y,corr_test_x,test_y)\n\n#chi2\nmodel(LassoCV(cv=cv, random_state=0,alphas=[1e-3, 1e-2, 1e-1, 1]),chi2_train_x,train_y,chi2_test_x,test_y)\n\n#mutual info \nmodel(LassoCV(cv=cv, random_state=0,alphas=[1e-3, 1e-2, 1e-1, 1]),mut_train_x,train_y,mut_test_x,test_y)\n\n#feature importance\nmodel(LassoCV(cv=cv, random_state=0,alphas=[1e-3, 1e-2, 1e-1, 1]),featimp_train_x,train_y,featimp_test_x,test_y)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:13.727160Z","iopub.execute_input":"2022-07-08T02:35:13.730719Z","iopub.status.idle":"2022-07-08T02:35:14.101242Z","shell.execute_reply.started":"2022-07-08T02:35:13.730650Z","shell.execute_reply":"2022-07-08T02:35:14.099993Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 4. ElasticNet","metadata":{}},{"cell_type":"code","source":"#raw dataset as well \nmodel(ElasticNet(random_state=0),train_x,train_y,test_x,test_y)\n\n#corr\nmodel(ElasticNet(random_state=0),corr_train_x,train_y,corr_test_x,test_y)\n\n#chi2\nmodel(ElasticNet(random_state=0),chi2_train_x,train_y,chi2_test_x,test_y)\n\n#mutual info \nmodel(ElasticNet(random_state=0),mut_train_x,train_y,mut_test_x,test_y)\n\n#feature importance\nmodel(ElasticNet(random_state=0),featimp_train_x,train_y,featimp_test_x,test_y)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:14.107014Z","iopub.execute_input":"2022-07-08T02:35:14.111017Z","iopub.status.idle":"2022-07-08T02:35:14.218034Z","shell.execute_reply.started":"2022-07-08T02:35:14.110950Z","shell.execute_reply":"2022-07-08T02:35:14.216676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 5. Random Forest Regressor\n\nFor this we will perform k fold cross validation first to try and get the best hyper-paramters. ","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import cross_val_score, GridSearchCV\n\ndef rfr_model(X, y,x_test,y_test):\n# Perform Grid-Search\n    cv=StratifiedKFold(n_splits=5, random_state=None, shuffle=False)\n    gsc = GridSearchCV(\n        estimator=RandomForestRegressor(),\n        param_grid={\n            'max_depth': range(3,7),\n            'n_estimators': (10, 50, 100, 1000),\n        },\n        cv=cv, scoring='neg_mean_squared_error', verbose=1,                         n_jobs=-1)\n    \n    grid_result = gsc.fit(X, y)\n    best_params = grid_result.best_params_\n    \n    rfr = RandomForestRegressor(max_depth=best_params[\"max_depth\"], n_estimators=best_params[\"n_estimators\"],                               random_state=False, verbose=False)\n    return model(rfr,X,y,x_test,y_test)\n\n\n\nrfr_model(train_x,train_y,test_x,test_y)\nrfr_model(corr_train_x,train_y,corr_test_x,test_y)\nrfr_model(chi2_train_x,train_y,chi2_test_x,test_y)\nrfr_model(mut_train_x,train_y,mut_test_x,test_y)\nrfr_model(featimp_train_x,train_y,featimp_test_x,test_y)\n\n\n\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:35:14.224796Z","iopub.execute_input":"2022-07-08T02:35:14.228368Z","iopub.status.idle":"2022-07-08T02:37:50.330943Z","shell.execute_reply.started":"2022-07-08T02:35:14.228295Z","shell.execute_reply":"2022-07-08T02:37:50.329523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"as our best model turns out to be random forest regressor with the feat imptortance vairable, let;s use that as the final model. ","metadata":{}},{"cell_type":"code","source":"cv=StratifiedKFold(n_splits=5, random_state=None, shuffle=False)\ngsc = GridSearchCV(\n        estimator=RandomForestRegressor(),\n        param_grid={\n            'max_depth': range(3,7),\n            'n_estimators': (10, 50, 100, 1000),\n        },\n        cv=cv, scoring='neg_mean_squared_error', verbose=1,                         n_jobs=-1)\n\ngrid_result = gsc.fit(featimp_train_x,train_y)\nbest_params = grid_result.best_params_\n    \nrfr = RandomForestRegressor(max_depth=best_params[\"max_depth\"], n_estimators=best_params[\"n_estimators\"],                               random_state=False, verbose=False)\nrfr.fit(featimp_train_x,train_y)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:48:40.585040Z","iopub.execute_input":"2022-07-08T02:48:40.585457Z","iopub.status.idle":"2022-07-08T02:49:11.997022Z","shell.execute_reply.started":"2022-07-08T02:48:40.585411Z","shell.execute_reply":"2022-07-08T02:49:11.995700Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we prep the test set before going in. ","metadata":{}},{"cell_type":"code","source":"testing=pd.read_csv('../input/house-prices-advanced-regression-techniques/test.csv')\ntesting.head()\ntest_id=testing['Id']\ntesting=testing.drop('Id',axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:59:38.119462Z","iopub.execute_input":"2022-07-08T02:59:38.119902Z","iopub.status.idle":"2022-07-08T02:59:38.146108Z","shell.execute_reply.started":"2022-07-08T02:59:38.119867Z","shell.execute_reply":"2022-07-08T02:59:38.144839Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"testing=testing[list(featimp_train_x.columns)]\nnumeric_data = testing.select_dtypes(include=[np.number])\ncategorical_data = testing.select_dtypes(exclude=[np.number])\n\n\ndef encode(dataset):\n    temp=[]\n    frequency=[]\n    for x in dataset.columns:\n    \n        if x.endswith('Qual'):\n            temp.append(x)\n        elif x.endswith('Cond'):\n            temp.append(x)\n\n        elif x.endswith('QC'):\n            temp.append(x)\n\n        else: \n            frequency.append(x)\n\n        \n    encoder=OrdinalEncoder()\n    encoder.fit(dataset[temp])\n    dataset[temp]=encoder.transform(dataset[temp])\n    \n    for y in frequency:\n        temp=dict(dataset[y].value_counts())\n        dataset[y]=dataset[y].replace(temp)\n    return dataset\n\n\ntrain_x_cat=encode(categorical_data)\n\ntrain_x_cat = train_x_cat.reset_index()\nnumeric_data=numeric_data.reset_index()\n\n\nsubmission=pd.concat([numeric_data,train_x_cat],axis=1)\nsubmission=submission.drop('index',axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:59:41.173552Z","iopub.execute_input":"2022-07-08T02:59:41.174135Z","iopub.status.idle":"2022-07-08T02:59:41.209131Z","shell.execute_reply.started":"2022-07-08T02:59:41.174084Z","shell.execute_reply":"2022-07-08T02:59:41.208292Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.isna().sum().sort_values()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:59:44.819149Z","iopub.execute_input":"2022-07-08T02:59:44.820175Z","iopub.status.idle":"2022-07-08T02:59:44.829550Z","shell.execute_reply.started":"2022-07-08T02:59:44.820122Z","shell.execute_reply":"2022-07-08T02:59:44.828418Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let;s impute by mode and median. ","metadata":{}},{"cell_type":"code","source":"submission['BsmtUnfSF']=submission['BsmtUnfSF'].fillna(submission['BsmtUnfSF'].median())\nsubmission['BsmtFinSF1']=submission['BsmtFinSF1'].fillna(submission['BsmtFinSF1'].median())\nsubmission['MasVnrArea']=submission['MasVnrArea'].fillna(submission['MasVnrArea'].median())\nsubmission['MSZoning']=submission['MSZoning'].fillna(submission['MSZoning'].mode()[0])","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:59:47.777431Z","iopub.execute_input":"2022-07-08T02:59:47.777841Z","iopub.status.idle":"2022-07-08T02:59:47.789636Z","shell.execute_reply.started":"2022-07-08T02:59:47.777809Z","shell.execute_reply":"2022-07-08T02:59:47.788706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.isna().sum().sort_values()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:59:50.072320Z","iopub.execute_input":"2022-07-08T02:59:50.073530Z","iopub.status.idle":"2022-07-08T02:59:50.083664Z","shell.execute_reply.started":"2022-07-08T02:59:50.073486Z","shell.execute_reply":"2022-07-08T02:59:50.082693Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred=rfr.predict(submission)\ny_pred[0:5]","metadata":{"execution":{"iopub.status.busy":"2022-07-08T02:59:53.228085Z","iopub.execute_input":"2022-07-08T02:59:53.229099Z","iopub.status.idle":"2022-07-08T02:59:53.257623Z","shell.execute_reply.started":"2022-07-08T02:59:53.229056Z","shell.execute_reply":"2022-07-08T02:59:53.256495Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.DataFrame()\nsubmission[\"Id\"] = test_id\nsubmission[\"SalePrice\"] = y_pred\nsubmission.to_csv(\"submission.csv\", index = False)\nsubmission.head(n = 10)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T03:00:05.509628Z","iopub.execute_input":"2022-07-08T03:00:05.509998Z","iopub.status.idle":"2022-07-08T03:00:05.533874Z","shell.execute_reply.started":"2022-07-08T03:00:05.509969Z","shell.execute_reply":"2022-07-08T03:00:05.532691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## COnclusion ","metadata":{}},{"cell_type":"markdown","source":"THis was a fun dataset and comp to work on, and i think some poitners to note of moving here. \n\n1. Perhaps one more extra step that could be taken is to further hyperparamter tune the random forest regressor so that we might get better results as a result \n\n2. Should we split the train into a frutehr trian and est here? Not sure if it makes much of a difference here. Shall try not splitting the next itme round to see what happens. But then we will have to rely on the cross cval score and it might be uncertain to wether those would be mroe beneficial as compareed to doign a frutehr train test split. \n\n3. What if we included mroe features? the reason i did not include so many is to prevent the case of the csurse of dimensionality from happening, as that may lead to futher reprecussions. \n\n\n\nTHank you and please give me your comments on my notebooks on things that i could imporve on or things that i mgith ave done wrongly! :) ","metadata":{}}]}