{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# The Ames Housing Price predict\n## Created By : Dwi Pamuji Bagaskara\nSource : https://www.kaggle.com/datasets/sakshigoyal7/credit-card-customers","metadata":{}},{"cell_type":"markdown","source":"# Content\n---\n\n1. Business Problem Understanding\n1. Data Understanding\n1. Exploratory Data Analysis\n1. Data Preprocessing\n1. Modeling and Evaluation\n1. Hyperparameter Tuning\n1. Compare Model Predict with Actual\n1. Test Machine Learning\n1. Conclution","metadata":{}},{"cell_type":"markdown","source":"# **Business Problem Understanding**\n---","metadata":{"toc-hr-collapsed":true}},{"cell_type":"markdown","source":"<font size = \"4\">**Context :**</font>\nHouse is one of the primary needs, as a customer One of the difficulties in owning a house is determining the price of the house based on the facilities provided. Is the price of the house in accordance with the existing facilities or is the price too expensive. For home sellers, determining the ideal house price for sale is also very important. that sellers don't lose their sales, or no one want to buys the house cause too expensive\n\n<font size = \"4\">**Problem Statement :**</font>\nDetermine the price of the house whether the price is appropriate\n\n<font size = \"4\">**Goals :**</font>\nhelp customers to predict the price of the house they want to buy or want to sell\n\n<font size = \"4\">**Analytic Approach :**</font>\nso what we need to do is analyze the data and find patterns between houses. then we will make machine learning regression for predict house price. **which will be useful for customers to determine the price of the house they want to buy or what they sell**\n\n<font size = \"4\">**Metric Evaluation**</font>\nEvaluation metrics to be used are RMSE and R-Squared, RMSE is the mean value of the square root of the error, means the model is more accurate in predicting the price according to the limitations of the features used. R-squared is used to determine how well the model can represent the overall variance of the data. The closer to 1, the more fit the model is to the observation data.\n","metadata":{}},{"cell_type":"markdown","source":"# Data Understanding\n---","metadata":{"toc-hr-collapsed":true}},{"cell_type":"markdown","source":"## **Data fields present in the dataset** ","metadata":{"toc-hr-collapsed":true}},{"cell_type":"markdown","source":"> ### *Categorical Features or Object Type Features*\n\nMSSubClass: Identifies the type of dwelling involved in the sale.\t\n\n        20\t1-STORY 1946 & NEWER ALL STYLES\n        30\t1-STORY 1945 & OLDER\n        40\t1-STORY W/FINISHED ATTIC ALL AGES\n        45\t1-1/2 STORY - UNFINISHED ALL AGES\n        50\t1-1/2 STORY FINISHED ALL AGES\n        60\t2-STORY 1946 & NEWER\n        70\t2-STORY 1945 & OLDER\n        75\t2-1/2 STORY ALL AGES\n        80\tSPLIT OR MULTI-LEVEL\n        85\tSPLIT FOYER\n        90\tDUPLEX - ALL STYLES AND AGES\n       120\t1-STORY PUD (Planned Unit Development) - 1946 & NEWER\n       150\t1-1/2 STORY PUD - ALL AGES\n       160\t2-STORY PUD - 1946 & NEWER\n       180\tPUD - MULTILEVEL - INCL SPLIT LEV/FOYER\n       190\t2 FAMILY CONVERSION - ALL STYLES AND AGES\n\nMSZoning: Identifies the general zoning classification of the sale.\n\t\t\n       A\tAgriculture\n       C\tCommercial\n       FV\tFloating Village Residential\n       I\tIndustrial\n       RH\tResidential High Density\n       RL\tResidential Low Density\n       RP\tResidential Low Density Park \n       RM\tResidential Medium Density\n       \nStreet: Type of road access to property\n\n       Grvl\tGravel\t\n       Pave\tPaved\n       \t\nAlley: Type of alley access to property\n\n       Grvl\tGravel\n       Pave\tPaved\n       NA \tNo alley access\n\t\t\nLotShape: General shape of property\n\n       Reg\tRegular\t\n       IR1\tSlightly irregular\n       IR2\tModerately Irregular\n       IR3\tIrregular\n       \nLandContour: Flatness of the property\n\n       Lvl\tNear Flat/Level\t\n       Bnk\tBanked - Quick and significant rise from street grade to building\n       HLS\tHillside - Significant slope from side to side\n       Low\tDepression\n\t\t\nUtilities: Type of utilities available\n\t\t\n       AllPub\tAll public Utilities (E,G,W,& S)\t\n       NoSewr\tElectricity, Gas, and Water (Septic Tank)\n       NoSeWa\tElectricity and Gas Only\n       ELO\tElectricity only\t\n\t\nLotConfig: Lot configuration\n\n       Inside\tInside lot\n       Corner\tCorner lot\n       CulDSac\tCul-de-sac\n       FR2\tFrontage on 2 sides of property\n       FR3\tFrontage on 3 sides of property\n\t\nLandSlope: Slope of property\n\t\t\n       Gtl\tGentle slope\n       Mod\tModerate Slope\t\n       Sev\tSevere Slope\n\t\nNeighborhood: Physical locations within Ames city limits\n\n       Blmngtn\tBloomington Heights\n       Blueste\tBluestem\n       BrDale\tBriardale\n       BrkSide\tBrookside\n       ClearCr\tClear Creek\n       CollgCr\tCollege Creek\n       Crawfor\tCrawford\n       Edwards\tEdwards\n       Gilbert\tGilbert\n       IDOTRR\tIowa DOT and Rail Road\n       MeadowV\tMeadow Village\n       Mitchel\tMitchell\n       Names\tNorth Ames\n       NoRidge\tNorthridge\n       NPkVill\tNorthpark Villa\n       NridgHt\tNorthridge Heights\n       NWAmes\tNorthwest Ames\n       OldTown\tOld Town\n       SWISU\tSouth & West of Iowa State University\n       Sawyer\tSawyer\n       SawyerW\tSawyer West\n       Somerst\tSomerset\n       StoneBr\tStone Brook\n       Timber\tTimberland\n       Veenker\tVeenker\n\t\t\t\nCondition1: Proximity to various conditions\n\t\n       Artery\tAdjacent to arterial street\n       Feedr\tAdjacent to feeder street\t\n       Norm\tNormal\t\n       RRNn\tWithin 200' of North-South Railroad\n       RRAn\tAdjacent to North-South Railroad\n       PosN\tNear positive off-site feature--park, greenbelt, etc.\n       PosA\tAdjacent to postive off-site feature\n       RRNe\tWithin 200' of East-West Railroad\n       RRAe\tAdjacent to East-West Railroad\n\t\nCondition2: Proximity to various conditions (if more than one is present)\n\t\t\n       Artery\tAdjacent to arterial street\n       Feedr\tAdjacent to feeder street\t\n       Norm\tNormal\t\n       RRNn\tWithin 200' of North-South Railroad\n       RRAn\tAdjacent to North-South Railroad\n       PosN\tNear positive off-site feature--park, greenbelt, etc.\n       PosA\tAdjacent to postive off-site feature\n       RRNe\tWithin 200' of East-West Railroad\n       RRAe\tAdjacent to East-West Railroad\n\t\nBldgType: Type of dwelling\n\t\t\n       1Fam\tSingle-family Detached\t\n       2FmCon\tTwo-family Conversion; originally built as one-family dwelling\n       Duplx\tDuplex\n       TwnhsE\tTownhouse End Unit\n       TwnhsI\tTownhouse Inside Unit\n\t\nHouseStyle: Style of dwelling\n\t\n       1Story\tOne story\n       1.5Fin\tOne and one-half story: 2nd level finished\n       1.5Unf\tOne and one-half story: 2nd level unfinished\n       2Story\tTwo story\n       2.5Fin\tTwo and one-half story: 2nd level finished\n       2.5Unf\tTwo and one-half story: 2nd level unfinished\n       SFoyer\tSplit Foyer\n       SLvl\tSplit Level\n\t\nOverallQual: Rates the overall material and finish of the house\n\n       10\tVery Excellent\n       9\tExcellent\n       8\tVery Good\n       7\tGood\n       6\tAbove Average\n       5\tAverage\n       4\tBelow Average\n       3\tFair\n       2\tPoor\n       1\tVery Poor\n\t\nOverallCond: Rates the overall condition of the house\n\n       10\tVery Excellent\n       9\tExcellent\n       8\tVery Good\n       7\tGood\n       6\tAbove Average\t\n       5\tAverage\n       4\tBelow Average\t\n       3\tFair\n       2\tPoor\n       1\tVery Poor\n\nRoofStyle: Type of roof\n\n       Flat\tFlat\n       Gable\tGable\n       Gambrel\tGabrel (Barn)\n       Hip\tHip\n       Mansard\tMansard\n       Shed\tShed\n\t\t\nRoofMatl: Roof material\n\n       ClyTile\tClay or Tile\n       CompShg\tStandard (Composite) Shingle\n       Membran\tMembrane\n       Metal\tMetal\n       Roll\tRoll\n       Tar&Grv\tGravel & Tar\n       WdShake\tWood Shakes\n       WdShngl\tWood Shingles\n\t\t\nExterior1st: Exterior covering on house\n\n       AsbShng\tAsbestos Shingles\n       AsphShn\tAsphalt Shingles\n       BrkComm\tBrick Common\n       BrkFace\tBrick Face\n       CBlock\tCinder Block\n       CemntBd\tCement Board\n       HdBoard\tHard Board\n       ImStucc\tImitation Stucco\n       MetalSd\tMetal Siding\n       Other\tOther\n       Plywood\tPlywood\n       PreCast\tPreCast\t\n       Stone\tStone\n       Stucco\tStucco\n       VinylSd\tVinyl Siding\n       Wd Sdng\tWood Siding\n       WdShing\tWood Shingles\n\t\nExterior2nd: Exterior covering on house (if more than one material)\n\n       AsbShng\tAsbestos Shingles\n       AsphShn\tAsphalt Shingles\n       BrkComm\tBrick Common\n       BrkFace\tBrick Face\n       CBlock\tCinder Block\n       CemntBd\tCement Board\n       HdBoard\tHard Board\n       ImStucc\tImitation Stucco\n       MetalSd\tMetal Siding\n       Other\tOther\n       Plywood\tPlywood\n       PreCast\tPreCast\n       Stone\tStone\n       Stucco\tStucco\n       VinylSd\tVinyl Siding\n       Wd Sdng\tWood Siding\n       WdShing\tWood Shingles\n\t\nMasVnrType: Masonry veneer type\n\n       BrkCmn\tBrick Common\n       BrkFace\tBrick Face\n       CBlock\tCinder Block\n       NA\t    None\n       Stone\tStone\n       \nExterQual: Evaluates the quality of the material on the exterior \n\t\t\n       Ex\tExcellent\n       Gd\tGood\n       TA\tAverage/Typical\n       Fa\tFair\n       Po\tPoor\n\t\t\nExterCond: Evaluates the present condition of the material on the exterior\n\t\t\n       Ex\tExcellent\n       Gd\tGood\n       TA\tAverage/Typical\n       Fa\tFair\n       Po\tPoor\n\t\t\nFoundation: Type of foundation\n\t\t\n       BrkTil\tBrick & Tile\n       CBlock\tCinder Block\n       PConc\tPoured Contrete\t\n       Slab\tSlab\n       Stone\tStone\n       Wood\tWood\n\t\t\nBsmtQual: Evaluates the height of the basement\n\n       Ex\tExcellent (100+ inches)\t\n       Gd\tGood (90-99 inches)\n       TA\tTypical (80-89 inches)\n       Fa\tFair (70-79 inches)\n       Po\tPoor (<70 inches\n       NA\tNo Basement\n\t\t\nBsmtCond: Evaluates the general condition of the basement\n\n       Ex\tExcellent\n       Gd\tGood\n       TA\tTypical - slight dampness allowed\n       Fa\tFair - dampness or some cracking or settling\n       Po\tPoor - Severe cracking, settling, or wetness\n       NA\tNo Basement\n\t\nBsmtExposure: Refers to walkout or garden level walls\n\n       Gd\tGood Exposure\n       Av\tAverage Exposure (split levels or foyers typically score average or above)\t\n       Mn\tMimimum Exposure\n       No\tNo Exposure\n       NA\tNo Basement\n\t\nBsmtFinType1: Rating of basement finished area\n\n       GLQ\tGood Living Quarters\n       ALQ\tAverage Living Quarters\n       BLQ\tBelow Average Living Quarters\t\n       Rec\tAverage Rec Room\n       LwQ\tLow Quality\n       Unf\tUnfinshed\n       NA\tNo Basement\n       \nBsmtFinType2: Rating of basement finished area (if multiple types)\n\n       GLQ\tGood Living Quarters\n       ALQ\tAverage Living Quarters\n       BLQ\tBelow Average Living Quarters\t\n       Rec\tAverage Rec Room\n       LwQ\tLow Quality\n       Unf\tUnfinshed\n       NA\tNo Basement\n\nHeating: Type of heating\n\t\t\n       Floor\tFloor Furnace\n       GasA\tGas forced warm air furnace\n       GasW\tGas hot water or steam heat\n       Grav\tGravity furnace\t\n       OthW\tHot water or steam heat other than gas\n       Wall\tWall furnace\n\t\t\nHeatingQC: Heating quality and condition\n\n       Ex\tExcellent\n       Gd\tGood\n       TA\tAverage/Typical\n       Fa\tFair\n       Po\tPoor\n\nCentralAir: Central air conditioning\n\n       N\tNo\n       Y\tYes\n\t\t\nElectrical: Electrical system\n\n       SBrkr\tStandard Circuit Breakers & Romex\n       FuseA\tFuse Box over 60 AMP and all Romex wiring (Average)\t\n       FuseF\t60 AMP Fuse Box and mostly Romex wiring (Fair)\n       FuseP\t60 AMP Fuse Box and mostly knob & tube wiring (poor)\n       Mix\tMixed\n       \nKitchenQual: Kitchen quality\n\n       Ex\tExcellent\n       Gd\tGood\n       TA\tTypical/Average\n       Fa\tFair\n       Po\tPoor\n\nFunctional: Home functionality (Assume typical unless deductions are warranted)\n\n       Typ\tTypical Functionality\n       Min1\tMinor Deductions 1\n       Min2\tMinor Deductions 2\n       Mod\tModerate Deductions\n       Maj1\tMajor Deductions 1\n       Maj2\tMajor Deductions 2\n       Sev\tSeverely Damaged\n       Sal\tSalvage only\n       \nFireplaceQu: Fireplace quality\n\n       Ex\tExcellent - Exceptional Masonry Fireplace\n       Gd\tGood - Masonry Fireplace in main level\n       TA\tAverage - Prefabricated Fireplace in main living area or Masonry Fireplace in basement\n       Fa\tFair - Prefabricated Fireplace in basement\n       Po\tPoor - Ben Franklin Stove\n       NA\tNo Fireplace\n\t\t\nGarageType: Garage location\n\t\t\n       2Types\tMore than one type of garage\n       Attchd\tAttached to home\n       Basment\tBasement Garage\n       BuiltIn\tBuilt-In (Garage part of house - typically has room above garage)\n       CarPort\tCar Port\n       Detchd\tDetached from home\n       NA\tNo Garage\n\nGarageFinish: Interior finish of the garage\n\n       Fin\tFinished\n       RFn\tRough Finished\t\n       Unf\tUnfinished\n       NA\tNo Garage\n\nGarageQual: Garage quality\n\n       Ex\tExcellent\n       Gd\tGood\n       TA\tTypical/Average\n       Fa\tFair\n       Po\tPoor\n       NA\tNo Garage\n\t\t\nGarageCond: Garage condition\n\n       Ex\tExcellent\n       Gd\tGood\n       TA\tTypical/Average\n       Fa\tFair\n       Po\tPoor\n       NA\tNo Garage\n\t\t\nPavedDrive: Paved driveway\n\n       Y\tPaved \n       P\tPartial Pavement\n       N\tDirt/Gravel\n\nPoolQC: Pool quality\n\t\t\n       Ex\tExcellent\n       Gd\tGood\n       TA\tAverage/Typical\n       Fa\tFair\n       NA\tNo Pool\n\t\t\nFence: Fence quality\n\t\t\n       GdPrv\tGood Privacy\n       MnPrv\tMinimum Privacy\n       GdWo\tGood Wood\n       MnWw\tMinimum Wood/Wire\n       NA\tNo Fence\n\t\nMiscFeature: Miscellaneous feature not covered in other categories\n\t\t\n       Elev\tElevator\n       Gar2\t2nd Garage (if not described in garage section)\n       Othr\tOther\n       Shed\tShed (over 100 SF)\n       TenC\tTennis Court\n       NA\tNone\n\nSaleType: Type of sale\n\t\t\n       WD \tWarranty Deed - Conventional\n       CWD\tWarranty Deed - Cash\n       VWD\tWarranty Deed - VA Loan\n       New\tHome just constructed and sold\n       COD\tCourt Officer Deed/Estate\n       Con\tContract 15% Down payment regular terms\n       ConLw\tContract Low Down payment and low interest\n       ConLI\tContract Low Interest\n       ConLD\tContract Low Down\n       Oth\tOther\n\t\t\nSaleCondition: Condition of sale\n\n       Normal\tNormal Sale\n       Abnorml\tAbnormal Sale -  trade, foreclosure, short sale\n       AdjLand\tAdjoining Land Purchase\n       Alloca\tAllocation - two linked properties with separate deeds, typically condo with a garage unit\t\n       Family\tSale between family members\n       Partial\tHome was not completed when last assessed (associated with New Homes)\n       \n> ### *Numerical Features or Numeric data Type Features*\n\nLotFrontage: Linear feet of street connected to property\n\nLotArea: Lot size in square feet\n\nYearBuilt: Original construction date\n\nYearRemodAdd: Remodel date (same as construction date if no remodeling or additions)\n\nMasVnrArea: Masonry veneer area in square feet\n\nBsmtFinSF1: Type 1 finished square feet\n\nBsmtFinSF2: Type 2 finished square feet\n\nBsmtUnfSF: Unfinished square feet of basement area\n\nTotalBsmtSF: Total square feet of basement area\n\n1stFlrSF: First Floor square feet\n \n2ndFlrSF: Second floor square feet\n\nLowQualFinSF: Low quality finished square feet (all floors)\n\nGrLivArea: Above grade (ground) living area square feet\n\nBsmtFullBath: Basement full bathrooms\n\nBsmtHalfBath: Basement half bathrooms\n\nFullBath: Full bathrooms above grade\n\nHalfBath: Half baths above grade\n\nBedroom: Bedrooms above grade (does NOT include basement bedrooms)\n\nKitchen: Kitchens above grade\n\nTotRmsAbvGrd: Total rooms above grade (does not include bathrooms) \n\nFireplaces: Number of fireplaces\n\nGarageYrBlt: Year garage was built\n\nGarageCars: Size of garage in car capacity\n\nGarageArea: Size of garage in square feet\n\nWoodDeckSF: Wood deck area in square feet\n\nOpenPorchSF: Open porch area in square feet\n\nEnclosedPorch: Enclosed porch area in square feet\n\n3SsnPorch: Three season porch area in square feet\n\nScreenPorch: Screen porch area in square feet\n\nPoolArea: Pool area in square feet\n\nMiscVal: $Value of miscellaneous feature\n\nMoSold: Month Sold (MM)\n\nYrSold: Year Sold (YYYY)","metadata":{"tags":[]}},{"cell_type":"code","source":"# Dataframe\nimport pandas as pd\nimport numpy as np\n\n# Data Visualitation\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n#Missing Value\nimport missingno\n\n# Handling Warning\nimport warnings\nwarnings.filterwarnings('ignore')\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:03.490191Z","iopub.execute_input":"2022-08-08T13:04:03.490533Z","iopub.status.idle":"2022-08-08T13:04:03.505183Z","shell.execute_reply.started":"2022-08-08T13:04:03.490507Z","shell.execute_reply":"2022-08-08T13:04:03.504004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train = pd.read_csv(\"/kaggle/input/house-prices-advanced-regression-techniques/train.csv\")\ndf_test = pd.read_csv(\"/kaggle/input/house-prices-advanced-regression-techniques/test.csv\")\n\n# We will concat both dataframe for EDA and drop Column SalePrice at df_train\ndf = pd.concat([df_train, df_test], axis = 0)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:03.507391Z","iopub.execute_input":"2022-08-08T13:04:03.508453Z","iopub.status.idle":"2022-08-08T13:04:03.586626Z","shell.execute_reply.started":"2022-08-08T13:04:03.508416Z","shell.execute_reply":"2022-08-08T13:04:03.585960Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.drop(\"SalePrice\", axis = 1, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:03.587948Z","iopub.execute_input":"2022-08-08T13:04:03.589080Z","iopub.status.idle":"2022-08-08T13:04:03.601962Z","shell.execute_reply.started":"2022-08-08T13:04:03.589029Z","shell.execute_reply":"2022-08-08T13:04:03.600724Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:03.605215Z","iopub.execute_input":"2022-08-08T13:04:03.606062Z","iopub.status.idle":"2022-08-08T13:04:03.635056Z","shell.execute_reply.started":"2022-08-08T13:04:03.606024Z","shell.execute_reply":"2022-08-08T13:04:03.633957Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"listItem = []\nfor col in df.columns :\n    listItem.append([col, df[col].dtype, df[col].isna().sum(), round((df[col].isna().sum()/len(df[col])) * 100,2),\n                    df[col].nunique(), list(df[col].drop_duplicates().sample(2).values)]);\n\ndfDesc = pd.DataFrame(columns=['dataFeatures', 'dataType', 'null', 'nullPct', 'unique', 'uniqueSample'],\n                     data=listItem)\ndfDesc","metadata":{"tags":[],"execution":{"iopub.status.busy":"2022-08-08T13:04:03.637632Z","iopub.execute_input":"2022-08-08T13:04:03.638987Z","iopub.status.idle":"2022-08-08T13:04:03.737178Z","shell.execute_reply.started":"2022-08-08T13:04:03.638949Z","shell.execute_reply":"2022-08-08T13:04:03.736063Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Exploratory Data Analysis (EDA)\n-----","metadata":{"tags":[],"toc-hr-collapsed":true}},{"cell_type":"markdown","source":"## EDA General\n> EDA Both of data (Data Train and Data Test)","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(10,8))\ntotal = float(len(df))\nax = sns.countplot(x=df[\"MSZoning\"], edgecolor = 'black', order = df[\"MSZoning\"].value_counts().index)\nplt.title('Persentage House Zoning', fontsize=20)\nplt.xlabel('House Zone', fontsize = 20)\nplt.xticks(size = 13)\nfor p in ax.patches:\n    percentage = '{:.2f}%'.format(100 * p.get_height()/total)\n    x = p.get_x() + p.get_width()\n    y = p.get_height()\n    ax.annotate(percentage, (x, y), ha='right', va='bottom', fontsize = 20)\n\n# plt.savefig('EDA 1.jpg')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:03.738356Z","iopub.execute_input":"2022-08-08T13:04:03.738640Z","iopub.status.idle":"2022-08-08T13:04:03.936173Z","shell.execute_reply.started":"2022-08-08T13:04:03.738617Z","shell.execute_reply":"2022-08-08T13:04:03.935146Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Resident Low Density","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(12,10))\ntotal = float(len(df))\nax = sns.countplot(x=df[\"SaleType\"], edgecolor = 'black', order = df[\"SaleType\"].value_counts().index)\nplt.title('Persentage Sale Type', fontsize=20)\nplt.xlabel('Sale Type', fontsize = 20)\nplt.xticks(size = 13)\nfor p in ax.patches:\n    percentage = '{:.2f}%'.format(100 * p.get_height()/total)\n    x = p.get_x() + p.get_width()\n    y = p.get_height()\n    ax.annotate(percentage, (x, y), ha='right', va='bottom', fontsize = 20)\n\n# plt.savefig('EDA 1.jpg')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:03.937570Z","iopub.execute_input":"2022-08-08T13:04:03.937944Z","iopub.status.idle":"2022-08-08T13:04:04.146093Z","shell.execute_reply.started":"2022-08-08T13:04:03.937915Z","shell.execute_reply":"2022-08-08T13:04:04.145029Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Mostly house sale with normal Condition","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(10,8))\ntotal = float(len(df))\nax = sns.countplot(x=df[\"SaleCondition\"], edgecolor = 'black', order = df[\"SaleCondition\"].value_counts().index)\nplt.title('Persentage Sale Condition ', fontsize=20)\nplt.xlabel('Sale Condition', fontsize = 20)\nplt.xticks(size = 13)\nfor p in ax.patches:\n    percentage = '{:.2f}%'.format(100 * p.get_height()/total)\n    x = p.get_x() + p.get_width()\n    y = p.get_height()\n    ax.annotate(percentage, (x, y), ha='right', va='bottom', fontsize = 20)\n\n# plt.savefig('EDA 1.jpg')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:04.147423Z","iopub.execute_input":"2022-08-08T13:04:04.147707Z","iopub.status.idle":"2022-08-08T13:04:04.323380Z","shell.execute_reply.started":"2022-08-08T13:04:04.147681Z","shell.execute_reply":"2022-08-08T13:04:04.322357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Mostly house sale with normal Condition","metadata":{}},{"cell_type":"markdown","source":"## House facilities","metadata":{}},{"cell_type":"code","source":"## Countplot for CentralAir and BedroomAbvGr\n\nplt.figure(figsize=(15,9))\nsns.countplot(x='BedroomAbvGr',hue='FullBath',palette='terrain',data=df).set(title=\"FullBath vs Bedroom\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:04.325748Z","iopub.execute_input":"2022-08-08T13:04:04.326080Z","iopub.status.idle":"2022-08-08T13:04:04.594518Z","shell.execute_reply.started":"2022-08-08T13:04:04.326054Z","shell.execute_reply":"2022-08-08T13:04:04.593319Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"most houses have 3 bedrooms and 2 Full Bathrooms","metadata":{}},{"cell_type":"markdown","source":"## EDA Data Train\n> EDA only use data train","metadata":{}},{"cell_type":"code","source":"max_saleprice = df_train.groupby('MSZoning')[['SalePrice']].max().reset_index().sort_values('SalePrice', ascending=False)\n","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:04.598501Z","iopub.execute_input":"2022-08-08T13:04:04.598754Z","iopub.status.idle":"2022-08-08T13:04:04.611063Z","shell.execute_reply.started":"2022-08-08T13:04:04.598731Z","shell.execute_reply":"2022-08-08T13:04:04.610362Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"max_saleprice_map = {\n                    \"A\" : \"Agriculture\",\n                    \"C (all)\"\t: \"Commercial\",\n                    \"FV\" :\t\"Floating Village Residential\",\n                    \"I\" :\t\"Industrial\",\n                    \"RH\" :\t\"Residential High Density\",\n                    \"RL\" :\t\"Residential Low Density\",\n                    \"RP\" :\t\"Residential Low Density Park\",\n                    \"RM\" :\t\"Residential Medium Density\"\n                    }\n\nmax_saleprice.loc[:, 'MSZoning'] = max_saleprice['MSZoning'].map(max_saleprice_map)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:04.612252Z","iopub.execute_input":"2022-08-08T13:04:04.612913Z","iopub.status.idle":"2022-08-08T13:04:04.620327Z","shell.execute_reply.started":"2022-08-08T13:04:04.612867Z","shell.execute_reply":"2022-08-08T13:04:04.619571Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"max_saleprice","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:04.621541Z","iopub.execute_input":"2022-08-08T13:04:04.621909Z","iopub.status.idle":"2022-08-08T13:04:04.637102Z","shell.execute_reply.started":"2022-08-08T13:04:04.621863Z","shell.execute_reply":"2022-08-08T13:04:04.636053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This is an interesting thing, The most expansive house at the Residential Low Density zone","metadata":{}},{"cell_type":"code","source":"## Getting Numeric Features\nobject_feat =list(columns for columns in df_train.select_dtypes([object]).columns)\nnumeric_feat =list(columns for columns in df_train.select_dtypes([float, int]).columns)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:04.638563Z","iopub.execute_input":"2022-08-08T13:04:04.639210Z","iopub.status.idle":"2022-08-08T13:04:04.652841Z","shell.execute_reply.started":"2022-08-08T13:04:04.639181Z","shell.execute_reply":"2022-08-08T13:04:04.651829Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(object_feat)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:04.653787Z","iopub.execute_input":"2022-08-08T13:04:04.655179Z","iopub.status.idle":"2022-08-08T13:04:04.667853Z","shell.execute_reply.started":"2022-08-08T13:04:04.655116Z","shell.execute_reply":"2022-08-08T13:04:04.667008Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"object_feat[0:5]","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:04.669066Z","iopub.execute_input":"2022-08-08T13:04:04.669305Z","iopub.status.idle":"2022-08-08T13:04:04.684979Z","shell.execute_reply.started":"2022-08-08T13:04:04.669275Z","shell.execute_reply":"2022-08-08T13:04:04.684102Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(numeric_feat)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:04.686287Z","iopub.execute_input":"2022-08-08T13:04:04.686519Z","iopub.status.idle":"2022-08-08T13:04:04.698959Z","shell.execute_reply.started":"2022-08-08T13:04:04.686496Z","shell.execute_reply":"2022-08-08T13:04:04.697744Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"numeric_feat[0:5]","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:04.699954Z","iopub.execute_input":"2022-08-08T13:04:04.700958Z","iopub.status.idle":"2022-08-08T13:04:04.711826Z","shell.execute_reply.started":"2022-08-08T13:04:04.700931Z","shell.execute_reply":"2022-08-08T13:04:04.711167Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## barplot for saleprice and the data type with objects\n\nplt.figure(figsize=(28,160))\nplotnumber=1\nfor i in object_feat:\n    ax=plt.subplot(15,3,plotnumber)\n    sns.barplot(x=df_train[i],y=df_train.SalePrice,palette='Set2')\n    plt.xlabel(i, size = 20)\n    plt.xticks(rotation=70, size = 15)\n    plotnumber+=1\nplt.show() ","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:04.712815Z","iopub.execute_input":"2022-08-08T13:04:04.714090Z","iopub.status.idle":"2022-08-08T13:04:15.336605Z","shell.execute_reply.started":"2022-08-08T13:04:04.714037Z","shell.execute_reply":"2022-08-08T13:04:15.333587Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Distribution plot \n\nplt.figure(figsize=(28,80))\nplotnumber=1\nfor i in numeric_feat:\n    ax=plt.subplot(10,4,plotnumber)\n    sns.distplot(x=df_train[i], color = 'g', \n                 kde_kws={\"color\": \"r\", \"lw\": 3, \"label\": \"KDE\"})\n    plt.xlabel(i, size = 20)\n    plt.xticks(rotation=70)\n    plotnumber+=1\nplt.show()","metadata":{"tags":[],"execution":{"iopub.status.busy":"2022-08-08T13:04:15.338362Z","iopub.execute_input":"2022-08-08T13:04:15.338699Z","iopub.status.idle":"2022-08-08T13:04:22.270837Z","shell.execute_reply.started":"2022-08-08T13:04:15.338653Z","shell.execute_reply":"2022-08-08T13:04:22.269501Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Preprocessing\n---","metadata":{"toc-hr-collapsed":true}},{"cell_type":"markdown","source":"## **Indentify Outlier**","metadata":{}},{"cell_type":"code","source":"len(numeric_feat)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:22.272080Z","iopub.execute_input":"2022-08-08T13:04:22.272325Z","iopub.status.idle":"2022-08-08T13:04:22.279209Z","shell.execute_reply.started":"2022-08-08T13:04:22.272301Z","shell.execute_reply":"2022-08-08T13:04:22.277826Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(60,80))\nplotnumber=1\nfor i in numeric_feat:\n    plt.subplot(13,3,plotnumber)\n    sns.boxplot(df_train[i], color = \"g\")\n    plt.title(i, fontsize = 42)\n    plt.xticks(fontsize = 35)\n    plotnumber+=1\nplt.tight_layout()\n# plt.savefig('Data Preprocessing 2 - Box Plot Check Outlier.jpg')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:22.280782Z","iopub.execute_input":"2022-08-08T13:04:22.281574Z","iopub.status.idle":"2022-08-08T13:04:27.640948Z","shell.execute_reply.started":"2022-08-08T13:04:22.281535Z","shell.execute_reply":"2022-08-08T13:04:27.639046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# creating function to generate IQR, lower limit, and Upper limit\nLower_Limit = 0\nUpper_Limit = 0\ndef find_outlier(df, feature):\n    q1 = df[feature].quantile(0.25)\n    q2 = df[feature].quantile(0.50)\n    q3 = df[feature].quantile(0.75)\n    iqr = q3 - q1\n    limit = iqr*1.5\n    print(f'IQR: {iqr}')\n    global Lower_Limit, Upper_Limit\n    Lower_Limit = q1 - limit\n    if Lower_Limit < 0 :\n        Lower_Limit = 0\n    Upper_Limit = q3 + limit\n    print(f'Lower_Limit: {Lower_Limit}')\n    print(f'median: {q2}')\n    print(f'Upper_Limit: {Upper_Limit}')\n    print('_________________________')","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:27.642325Z","iopub.execute_input":"2022-08-08T13:04:27.642887Z","iopub.status.idle":"2022-08-08T13:04:27.651020Z","shell.execute_reply.started":"2022-08-08T13:04:27.642847Z","shell.execute_reply":"2022-08-08T13:04:27.649704Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"numeric_feat[1]","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:27.651993Z","iopub.execute_input":"2022-08-08T13:04:27.652516Z","iopub.status.idle":"2022-08-08T13:04:27.672722Z","shell.execute_reply.started":"2022-08-08T13:04:27.652489Z","shell.execute_reply":"2022-08-08T13:04:27.670993Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Upper_Limit","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:27.674154Z","iopub.execute_input":"2022-08-08T13:04:27.674786Z","iopub.status.idle":"2022-08-08T13:04:27.684537Z","shell.execute_reply.started":"2022-08-08T13:04:27.674757Z","shell.execute_reply":"2022-08-08T13:04:27.683387Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check IQR, upper limit, and lower limit for each feature\n\nfor i in range(len(numeric_feat)) :\n    print(f'Outlier_{i+1} : ' + numeric_feat[i])\n    find_outlier(df_train, numeric_feat[i])\n    exec(f'outlier_{i+1} = df_train[(df_train[numeric_feat[i]] > Upper_Limit) | (df_train[numeric_feat[i]] < Lower_Limit)]')","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:27.685742Z","iopub.execute_input":"2022-08-08T13:04:27.686312Z","iopub.status.idle":"2022-08-08T13:04:27.791008Z","shell.execute_reply.started":"2022-08-08T13:04:27.686286Z","shell.execute_reply":"2022-08-08T13:04:27.790330Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in range(1, 39) :\n    print(f'outlier_{i},')","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:27.792265Z","iopub.execute_input":"2022-08-08T13:04:27.792725Z","iopub.status.idle":"2022-08-08T13:04:27.798980Z","shell.execute_reply.started":"2022-08-08T13:04:27.792697Z","shell.execute_reply":"2022-08-08T13:04:27.797741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# outlier_all = []\n# for i in range(len(numeric_feat)) :\n#     exec(f'outlier_all.append(outlier_{i+1})')\n\nout_all = pd.concat([outlier_1,\n                    outlier_2,\n                    outlier_3,\n                    outlier_4,\n                    outlier_5,\n                    outlier_6,\n                    outlier_7,\n                    outlier_8,\n                    outlier_9,\n                    outlier_10,\n                    outlier_11,\n                    outlier_12,\n                    outlier_13,\n                    outlier_14,\n                    outlier_15,\n                    outlier_16,\n                    outlier_17,\n                    outlier_18,\n                    outlier_19,\n                    outlier_20,\n                    outlier_21,\n                    outlier_22,\n                    outlier_23,\n                    outlier_24,\n                    outlier_25,\n                    outlier_26,\n                    outlier_27,\n                    outlier_28,\n                    outlier_29,\n                    outlier_30,\n                    outlier_31,\n                    outlier_32,\n                    outlier_33,\n                    outlier_34,\n                    outlier_35,\n                    outlier_36,\n                    outlier_37,\n                    outlier_38], axis = 0)\nout_all.drop_duplicates(inplace=True)\nout_all","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:27.800377Z","iopub.execute_input":"2022-08-08T13:04:27.800841Z","iopub.status.idle":"2022-08-08T13:04:27.912280Z","shell.execute_reply.started":"2022-08-08T13:04:27.800813Z","shell.execute_reply":"2022-08-08T13:04:27.911711Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"per_outlier = len(out_all)/len(df_train)*100\nper_outlier","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:27.918495Z","iopub.execute_input":"2022-08-08T13:04:27.919060Z","iopub.status.idle":"2022-08-08T13:04:27.925644Z","shell.execute_reply.started":"2022-08-08T13:04:27.919032Z","shell.execute_reply":"2022-08-08T13:04:27.924677Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Total outlier is 897 outlier or 61.4% from 1460, we will keep outlier","metadata":{}},{"cell_type":"markdown","source":"## Identify Duplicated Data ","metadata":{}},{"cell_type":"code","source":"df_train[df_train.duplicated()]","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:27.926890Z","iopub.execute_input":"2022-08-08T13:04:27.927151Z","iopub.status.idle":"2022-08-08T13:04:27.959721Z","shell.execute_reply.started":"2022-08-08T13:04:27.927127Z","shell.execute_reply":"2022-08-08T13:04:27.958705Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#check whether there is any duplicate value\npd.DataFrame({'Duplicated Data' : df_train[df_train.duplicated()].sum()})","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:27.961275Z","iopub.execute_input":"2022-08-08T13:04:27.961624Z","iopub.status.idle":"2022-08-08T13:04:27.991208Z","shell.execute_reply.started":"2022-08-08T13:04:27.961589Z","shell.execute_reply":"2022-08-08T13:04:27.990265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In this dataset not have duplicated data","metadata":{}},{"cell_type":"markdown","source":"## Identify Missing Value","metadata":{}},{"cell_type":"code","source":"# check missing value\nmiss = pd.DataFrame({'Missing Value' : df_train.isna().sum()})\nmiss[miss['Missing Value'] > 0 ]","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:27.992602Z","iopub.execute_input":"2022-08-08T13:04:27.992931Z","iopub.status.idle":"2022-08-08T13:04:28.007358Z","shell.execute_reply.started":"2022-08-08T13:04:27.992898Z","shell.execute_reply":"2022-08-08T13:04:28.006089Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Missing value heatmap\nplt.figure(figsize = (15, 10))\nsns.heatmap(df_train.isnull(), cbar=False);\nplt.title('Heatmap Missing Value \\n', size =20)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:28.008642Z","iopub.execute_input":"2022-08-08T13:04:28.008943Z","iopub.status.idle":"2022-08-08T13:04:29.421080Z","shell.execute_reply.started":"2022-08-08T13:04:28.008918Z","shell.execute_reply":"2022-08-08T13:04:29.420191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"19 feature have missing value\n\nwe check describe data for check info NA\n","metadata":{}},{"cell_type":"markdown","source":"Example NA in Categorical Feature\n\nAlley : Type of alley access to property\n\n       NA \tNo alley access","metadata":{}},{"cell_type":"code","source":"df_train[df_train['Alley'].isna()].head(10)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:29.422141Z","iopub.execute_input":"2022-08-08T13:04:29.422369Z","iopub.status.idle":"2022-08-08T13:04:29.448338Z","shell.execute_reply.started":"2022-08-08T13:04:29.422345Z","shell.execute_reply":"2022-08-08T13:04:29.447708Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"we check all variable has NA in this cataegori and compare with NaN in missing value, python read NA in categorical variable as Missing Vale (NaN)","metadata":{}},{"cell_type":"markdown","source":"Categorical variable that has a value of NA :\n1. Alley : NA \tNo alley access\n1. BsmtQual : NA\tNo Basement\n1. BsmtCond : NA\tNo Basement\n1. BsmtExposure : NA\tNo Basement\n1. BsmtFinType1 : NA\tNo Basement\n1. BsmtFinType2 : NA\tNo Basement\n1. FireplaceQu : NA\tNo Fireplace\n1. GarageType : NA\tNo Garage\n1. GarageYrBlt : NA\tNo Garage\n1. GarageFinish : NA\tNo Garage\n1. GarageQual : NA\tNo Garage\n1. GarageCond : NA\tNo Garage\n1. PoolQC : NA\tNo Pool\n1. Fence : NA\tNo Fence\n1. MiscFeature : NA\tNone","metadata":{}},{"cell_type":"markdown","source":"## Handling Missing Value","metadata":{}},{"cell_type":"code","source":"from sklearn.impute import SimpleImputer","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:29.449392Z","iopub.execute_input":"2022-08-08T13:04:29.449804Z","iopub.status.idle":"2022-08-08T13:04:29.762080Z","shell.execute_reply.started":"2022-08-08T13:04:29.449777Z","shell.execute_reply":"2022-08-08T13:04:29.760968Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"feature ```LotFrontage``` we asumse NA is haouse without Frontage we will imput NA with 0\n\nfeature ```MasVnrType``` we asumse NA is haouse without Veneer we will imput NA with **None**\n\nfeature ```MasVnrArea``` we asumse NA is haouse without Veneer we will imput NA with 0\n\nfeature ```Elictrical``` we asumse NA is haouse with standard Elictrical or mosly at hause, we will imput with simpleimputer frequency\n\nfeature ```GarageYrBlt``` we asumse NA is haouse without garage, cause not have garage garage build must NA, cause this data type is int we will imput 0\n\nfor 14 other features we will imput NA with **NA**, cause we asumse python read categorical 'NA' as Missing Value","metadata":{}},{"cell_type":"markdown","source":"**Scheme for Handling Missing Value :**\n1. Simple Imputer with value 0 : LotFrontage, MasVnrArea, GarageYrBlt\n1. Simple Imputer with value None : MasVnrType\n1. Simple Imputer with value Modus : Electrical\n1. Simple Imputer with value NA : Alley, BsmtQual, BsmtCond, BsmtExposure, BsmtFinType1, BsmtFinType2, FireplaceQu, GarageType, GarageFinish, GarageQual, GarageCond, PoolQC, Fence, MiscFeature","metadata":{}},{"cell_type":"code","source":"s_imputer_1 = SimpleImputer(strategy='constant', fill_value=0)\ns_imputer_2 = SimpleImputer(strategy='constant', fill_value=\"None\")\ns_imputer_3 = SimpleImputer(strategy='most_frequent')\ns_imputer_4 = SimpleImputer(strategy='constant', fill_value=\"NA\")","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:29.763394Z","iopub.execute_input":"2022-08-08T13:04:29.763645Z","iopub.status.idle":"2022-08-08T13:04:29.768734Z","shell.execute_reply.started":"2022-08-08T13:04:29.763622Z","shell.execute_reply":"2022-08-08T13:04:29.767612Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# filling missing value with iterative imputer\nfeat_1 = ['LotFrontage', 'MasVnrArea', 'GarageYrBlt']\nfeat_2 = ['MasVnrType']\nfeat_3 = ['Electrical']\nfeat_4 = ['Alley', 'BsmtQual', 'BsmtCond', 'BsmtExposure', 'BsmtFinType1', \n          'BsmtFinType2', 'FireplaceQu', 'GarageType', 'GarageFinish', \n          'GarageQual', 'GarageCond', 'PoolQC', 'Fence', 'MiscFeature']","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:29.770097Z","iopub.execute_input":"2022-08-08T13:04:29.770598Z","iopub.status.idle":"2022-08-08T13:04:29.787150Z","shell.execute_reply.started":"2022-08-08T13:04:29.770565Z","shell.execute_reply":"2022-08-08T13:04:29.785898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train[feat_1] = s_imputer_1.fit_transform(df_train[feat_1])\ndf_train[feat_2] = s_imputer_2.fit_transform(df_train[feat_2])\ndf_train[feat_3] = s_imputer_3.fit_transform(df_train[feat_3])\ndf_train[feat_4] = s_imputer_4.fit_transform(df_train[feat_4])","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:29.788637Z","iopub.execute_input":"2022-08-08T13:04:29.789355Z","iopub.status.idle":"2022-08-08T13:04:29.814985Z","shell.execute_reply.started":"2022-08-08T13:04:29.789320Z","shell.execute_reply":"2022-08-08T13:04:29.814256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check missing value\nmiss = pd.DataFrame({'Missing Value' : df_train.isna().sum()})\nmiss[miss['Missing Value'] > 0 ]","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:29.815957Z","iopub.execute_input":"2022-08-08T13:04:29.816291Z","iopub.status.idle":"2022-08-08T13:04:29.828044Z","shell.execute_reply.started":"2022-08-08T13:04:29.816267Z","shell.execute_reply":"2022-08-08T13:04:29.827049Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Missing value heatmap\nplt.figure(figsize = (8,6))\nsns.heatmap(df_train.isnull(), cbar=False);\nplt.title('Heatmap Missing Value \\n', size =20)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:29.830068Z","iopub.execute_input":"2022-08-08T13:04:29.831119Z","iopub.status.idle":"2022-08-08T13:04:30.450652Z","shell.execute_reply.started":"2022-08-08T13:04:29.831081Z","shell.execute_reply":"2022-08-08T13:04:30.449162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After Handling misisng value we see not missing value in dataset set","metadata":{}},{"cell_type":"markdown","source":"## Summary Data Train","metadata":{}},{"cell_type":"code","source":"# central tendency: mean, median\nmean = pd.DataFrame(df_train[numeric_feat].apply(np.mean)).T\nmedian = pd.DataFrame(df_train[numeric_feat].apply(np.median)).T\n\n# distribution: ,std, min, max, range, skew, kurtosis\nstd = pd.DataFrame(df_train[numeric_feat].apply(np.std)).T\nmin_value = pd.DataFrame(df_train[numeric_feat].apply(min)).T\nmax_value = pd.DataFrame(df_train[numeric_feat].apply(max)).T\nrange_value = pd.DataFrame(df_train[numeric_feat].apply(lambda x: x.max() - x.min())).T\nskewness = pd.DataFrame(df_train[numeric_feat].apply(lambda x: x.skew())).T\nkurtosis = pd.DataFrame(df_train[numeric_feat].apply(lambda x: x.kurtosis())).T\n\n# concatenates\nsummary_stats = pd.concat([min_value, max_value, range_value, mean, median, std, skewness, kurtosis]).T.reset_index()\nsummary_stats.columns = ['attributes','min','max', 'range','mean','median', 'std','skewness','kurtosis']","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:30.451883Z","iopub.execute_input":"2022-08-08T13:04:30.452197Z","iopub.status.idle":"2022-08-08T13:04:30.517123Z","shell.execute_reply.started":"2022-08-08T13:04:30.452169Z","shell.execute_reply":"2022-08-08T13:04:30.516003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"summary_stats","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:30.518573Z","iopub.execute_input":"2022-08-08T13:04:30.518923Z","iopub.status.idle":"2022-08-08T13:04:30.547591Z","shell.execute_reply.started":"2022-08-08T13:04:30.518886Z","shell.execute_reply":"2022-08-08T13:04:30.546339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"kurtosis: the sharpness of the peak of a frequency-distribution curve. the bigger number are sharper.\n* (kurtosis = 3) mezokurtic (ideal sharpness)  \n* (kurtosis < 3) platycrutic (flatter curve)  \n* (kurtosis > 3) leptokrutic (very sharp)  \n* (kurtosis = 0) flat\n* (kurtosis < 0) U shape\n\nskewness: the asymmetry of a distribution\n* (skewness = 0) normal distributed\n* (skewness ~ -1 ) negative skew(left skewed)\n* (skewness ~ 1) positive skew(right skewed)\n* (skewness <> [-1,1]) very skewed distribution","metadata":{}},{"cell_type":"markdown","source":"# **Modeling and Evaluation**\n---","metadata":{"tags":[],"toc-hr-collapsed":true}},{"cell_type":"code","source":"# Import library untuk modeling\n\nfrom sklearn.model_selection import train_test_split, cross_val_score, RandomizedSearchCV, GridSearchCV, KFold, StratifiedKFold\n\nimport category_encoders as ce\nfrom sklearn.preprocessing import OneHotEncoder, StandardScaler, LabelEncoder, RobustScaler, MinMaxScaler\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.pipeline import Pipeline\nfrom category_encoders import BinaryEncoder\nfrom sklearn.linear_model import LinearRegression, Lasso, Ridge\nfrom sklearn.neighbors import KNeighborsRegressor\nfrom sklearn.tree import DecisionTreeRegressor\nfrom sklearn.ensemble import RandomForestRegressor, AdaBoostRegressor, GradientBoostingRegressor\nfrom xgboost.sklearn import XGBRegressor\nfrom sklearn.compose import TransformedTargetRegressor\n\nfrom sklearn.metrics import r2_score, mean_squared_error, mean_absolute_error, mean_absolute_percentage_error, f1_score","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:30.548885Z","iopub.execute_input":"2022-08-08T13:04:30.549213Z","iopub.status.idle":"2022-08-08T13:04:31.063101Z","shell.execute_reply.started":"2022-08-08T13:04:30.549179Z","shell.execute_reply":"2022-08-08T13:04:31.061610Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Make copt Data Train\n\ndf_train_b_modeling = df_train.copy()","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:31.064653Z","iopub.execute_input":"2022-08-08T13:04:31.065024Z","iopub.status.idle":"2022-08-08T13:04:31.071965Z","shell.execute_reply.started":"2022-08-08T13:04:31.064996Z","shell.execute_reply":"2022-08-08T13:04:31.070787Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Approaching Categorical Features**","metadata":{}},{"cell_type":"markdown","source":"for this case we will use label encoder","metadata":{}},{"cell_type":"code","source":"le = LabelEncoder()\n\nfor i in object_feat:\n    df_train[i] = le.fit_transform(df_train[i].astype(str))\n\nprint (df_train.info())","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:31.073822Z","iopub.execute_input":"2022-08-08T13:04:31.074150Z","iopub.status.idle":"2022-08-08T13:04:31.139743Z","shell.execute_reply.started":"2022-08-08T13:04:31.074113Z","shell.execute_reply":"2022-08-08T13:04:31.138763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:31.140847Z","iopub.execute_input":"2022-08-08T13:04:31.141201Z","iopub.status.idle":"2022-08-08T13:04:31.161365Z","shell.execute_reply.started":"2022-08-08T13:04:31.141175Z","shell.execute_reply":"2022-08-08T13:04:31.160175Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Splitting Data Train","metadata":{}},{"cell_type":"code","source":"X = df_train.drop(['Id', 'SalePrice'], axis = 1) # All Features\ny = df_train['SalePrice'] # Target","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:31.162674Z","iopub.execute_input":"2022-08-08T13:04:31.162943Z","iopub.status.idle":"2022-08-08T13:04:31.170479Z","shell.execute_reply.started":"2022-08-08T13:04:31.162918Z","shell.execute_reply":"2022-08-08T13:04:31.169315Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(\n                                    X, \n                                    y,\n                                    test_size = 0.3, \n                                    random_state = 2022)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:31.172271Z","iopub.execute_input":"2022-08-08T13:04:31.173340Z","iopub.status.idle":"2022-08-08T13:04:31.188471Z","shell.execute_reply.started":"2022-08-08T13:04:31.173296Z","shell.execute_reply":"2022-08-08T13:04:31.187691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train.head(5)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:31.189527Z","iopub.execute_input":"2022-08-08T13:04:31.189773Z","iopub.status.idle":"2022-08-08T13:04:31.213266Z","shell.execute_reply.started":"2022-08-08T13:04:31.189751Z","shell.execute_reply":"2022-08-08T13:04:31.212428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"le = LabelEncoder()\nscaler = RobustScaler()\n\ntransformer = ColumnTransformer([\n                ('scaler', scaler, numeric_feat[1:-2]) # for select numeric feat without 'id' and 'SalePrice'\n], remainder = \"passthrough\")","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:31.214316Z","iopub.execute_input":"2022-08-08T13:04:31.215384Z","iopub.status.idle":"2022-08-08T13:04:31.220438Z","shell.execute_reply.started":"2022-08-08T13:04:31.215335Z","shell.execute_reply":"2022-08-08T13:04:31.219629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"transformer.fit(X_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:31.221577Z","iopub.execute_input":"2022-08-08T13:04:31.222708Z","iopub.status.idle":"2022-08-08T13:04:31.249783Z","shell.execute_reply.started":"2022-08-08T13:04:31.222645Z","shell.execute_reply":"2022-08-08T13:04:31.248800Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Benchmark 9 Model with evaluation RMSE","metadata":{}},{"cell_type":"code","source":"# Define algoritma\nlr = LinearRegression()\nls = Lasso(random_state = 2022)\nrd = Ridge(random_state = 2022)\nknn = KNeighborsRegressor()\ndt = DecisionTreeRegressor(random_state=2022)\nrf = RandomForestRegressor(random_state=2022)\nab = AdaBoostRegressor(random_state = 2022)\ngbr = GradientBoostingRegressor(random_state = 2022)\nxgb = XGBRegressor(random_state=2022)\nlist_model = {'Linier Regression' : lr, 'Lasso' :  ls, 'Ridge' :  rd, 'KNN' :  knn,\n              'Decision Tree' :  dt, 'Random Forest' :  rf, 'AdaBoost' :  ab,\n              'GradientBoost' :  gbr, 'XGBoost' :  xgb}","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:31.250815Z","iopub.execute_input":"2022-08-08T13:04:31.251037Z","iopub.status.idle":"2022-08-08T13:04:31.258352Z","shell.execute_reply.started":"2022-08-08T13:04:31.251014Z","shell.execute_reply":"2022-08-08T13:04:31.257035Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"score_RMSE = []\nnilai_mean_RMSE = []\nnilai_std_RMSE = []\ncrossval = KFold(n_splits=5, shuffle=True, random_state=2022)\ndef model_eval_RMSE(model, metric):\n    for i in model :\n        estimator = Pipeline([\n            ('transformer', transformer),\n            ('model', list_model[i])\n        ])\n        model_cv = (-cross_val_score(estimator, X_train, y_train, cv = crossval, scoring = metric, error_score='raise'))\n        score_RMSE.append(model_cv)\n        nilai_mean_RMSE.append(model_cv.mean())\n        nilai_std_RMSE.append(model_cv.std())","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:31.259714Z","iopub.execute_input":"2022-08-08T13:04:31.260798Z","iopub.status.idle":"2022-08-08T13:04:31.273304Z","shell.execute_reply.started":"2022-08-08T13:04:31.260758Z","shell.execute_reply":"2022-08-08T13:04:31.272098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_eval_RMSE(list_model, 'neg_root_mean_squared_error')","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:31.274440Z","iopub.execute_input":"2022-08-08T13:04:31.275719Z","iopub.status.idle":"2022-08-08T13:04:43.834550Z","shell.execute_reply.started":"2022-08-08T13:04:31.275679Z","shell.execute_reply":"2022-08-08T13:04:43.833571Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_rmse = pd.DataFrame({\n    'model' : list_model.keys(),\n    'RMSE_Score' : score_RMSE,\n    'Mean_RMSE': nilai_mean_RMSE,\n    'Std_RMSE': nilai_std_RMSE\n})\nmodel_rmse","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:43.836060Z","iopub.execute_input":"2022-08-08T13:04:43.836506Z","iopub.status.idle":"2022-08-08T13:04:43.854549Z","shell.execute_reply.started":"2022-08-08T13:04:43.836478Z","shell.execute_reply":"2022-08-08T13:04:43.852853Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"best_b_tuning_train = model_rmse.sort_values('Mean_RMSE', ascending = True).head(2)\nbest_b_tuning_train","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:43.856281Z","iopub.execute_input":"2022-08-08T13:04:43.856621Z","iopub.status.idle":"2022-08-08T13:04:43.867692Z","shell.execute_reply.started":"2022-08-08T13:04:43.856595Z","shell.execute_reply":"2022-08-08T13:04:43.866825Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Benchmark best 2 model X_Test","metadata":{}},{"cell_type":"code","source":"gbr = GradientBoostingRegressor(random_state = 2022)\nxgb = XGBRegressor(random_state=2022)\n\nmodels = {\n    'XGBoost' : xgb,\n    'GradientBoost': gbr\n}\n\nscore_rmse = []\nscore_r2 = [] \n\n# Prediksi pada test set\nfor i in models:\n\n    model = Pipeline([\n        ('preprocessing', transformer),\n        ('model', models[i])\n        ])\n\n    model.fit(X_train, y_train)\n    exec(f'y_pred_b_{i} = model.predict(X_test)')\n    exec(f'score_rmse.append(np.sqrt(mean_squared_error(y_test, y_pred_b_{i})))')\n    exec(f'score_r2.append(r2_score(y_test, y_pred_b_{i}))')\n    exec(f'global y_pred_b_{i}')\n\nbest_b_tuning_test = pd.DataFrame({'model' : models.keys(),\n                                    'RMSE': score_rmse, \n                                    'R-Squared' : score_r2\n                                   })\nbest_b_tuning_test","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:43.868802Z","iopub.execute_input":"2022-08-08T13:04:43.869141Z","iopub.status.idle":"2022-08-08T13:04:44.978973Z","shell.execute_reply.started":"2022-08-08T13:04:43.869119Z","shell.execute_reply":"2022-08-08T13:04:44.978126Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the test results with test data, it can be seen that **GradientBoost** outperformed **XGBoost** in all metrics. **but to prove it I will do hyperparameter tuning to both moddels to make sure which model has the best performance**","metadata":{}},{"cell_type":"markdown","source":"# **Hyperparameter Tuning**\n---","metadata":{"tags":[],"toc-hr-collapsed":true}},{"cell_type":"markdown","source":"## **Hyperparameter Tuning XGBoost**","metadata":{}},{"cell_type":"code","source":"# The maximum depth limits the number of nodes in the tree. \nmax_depth = list(np.arange(15, 31))\n\n# Learning rate\nlearning_rate = list(np.arange(1, 100)/100)\n\n# The number of boosting stages to perform\nn_estimators = list(np.arange(100, 251))\n\n# The fraction of samples to be used for fitting the individual base learners.\nsubsample = list(np.arange(2, 10)/10)\n\n# Gamma (min_impurity_decrease)\ngamma = list(np.arange(1, 11)) # Semakin besar nilainya, semakin konservatif/simpel modelnya\n\n# Number of features used for each tree (% of the total train set column)\ncolsample_bytree = list(np.arange(1, 11/10))\n\n# Alpha (regularization)\nreg_alpha = list(np.logspace(-3, 1, 10)) # Semakin besar nilainya, semakin konservatif/simpel modelnya\n\n# Hyperparam space XGboost\nhyperparam_space_xgb = {\n    'model__max_depth': max_depth, \n    'model__learning_rate': learning_rate,\n    'model__n_estimators': n_estimators,\n    'model__subsample': subsample,\n    'model__gamma': gamma,\n    'model__colsample_bytree': colsample_bytree,\n    'model__reg_alpha': reg_alpha\n}","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:44.980066Z","iopub.execute_input":"2022-08-08T13:04:44.980292Z","iopub.status.idle":"2022-08-08T13:04:44.987194Z","shell.execute_reply.started":"2022-08-08T13:04:44.980268Z","shell.execute_reply":"2022-08-08T13:04:44.986303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Benchmark model dengan hyperparameter tuning\nxgb = XGBRegressor(random_state=2022)\n\n# Membuat algorithm chains\nestimator_xgb = Pipeline([\n        ('preprocessing', transformer),\n        ('model', xgb)\n        ])\n\ncrossval = KFold(n_splits=5, shuffle=True, random_state=1)\n\n# Hyperparameter tuning\nrandom_xgb = RandomizedSearchCV(\n    estimator_xgb, \n    param_distributions = hyperparam_space_xgb,\n    n_iter = 50,\n    cv = crossval, \n    scoring = ['neg_root_mean_squared_error', 'r2'], \n    n_jobs = -1,\n    refit = 'neg_root_mean_squared_error', # Hanya bisa memilih salah stau metric untuk optimisasi\n    random_state = 2022 \n)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:44.988282Z","iopub.execute_input":"2022-08-08T13:04:44.988755Z","iopub.status.idle":"2022-08-08T13:04:45.001001Z","shell.execute_reply.started":"2022-08-08T13:04:44.988728Z","shell.execute_reply":"2022-08-08T13:04:45.000226Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"random_xgb.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:04:45.001949Z","iopub.execute_input":"2022-08-08T13:04:45.002930Z","iopub.status.idle":"2022-08-08T13:08:35.231526Z","shell.execute_reply.started":"2022-08-08T13:04:45.002895Z","shell.execute_reply":"2022-08-08T13:08:35.228888Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.DataFrame(random_xgb.cv_results_).sort_values(by = ['rank_test_neg_root_mean_squared_error']).head(5)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:08:35.232891Z","iopub.execute_input":"2022-08-08T13:08:35.233211Z","iopub.status.idle":"2022-08-08T13:08:35.279709Z","shell.execute_reply.started":"2022-08-08T13:08:35.233179Z","shell.execute_reply":"2022-08-08T13:08:35.278031Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('XGBoost')\nprint('Best_score:', random_xgb.best_score_*-1)\nprint('Best_params:', random_xgb.best_params_)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:08:35.280549Z","iopub.execute_input":"2022-08-08T13:08:35.280809Z","iopub.status.idle":"2022-08-08T13:08:35.287835Z","shell.execute_reply.started":"2022-08-08T13:08:35.280783Z","shell.execute_reply":"2022-08-08T13:08:35.286273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_xgb_a_tuning = random_xgb.best_estimator_\ncrossval = KFold(n_splits=5, shuffle=True, random_state=2022)\n\nscore_RMSE = []\nnilai_mean_RMSE = []\nnilai_std_RMSE = []\n\nmodel_cv = (-cross_val_score(model_xgb_a_tuning, X_train, y_train, cv = crossval, scoring = 'neg_root_mean_squared_error'))\nscore_RMSE.append(model_cv)\nnilai_mean_RMSE.append(model_cv.mean())\nnilai_std_RMSE.append(model_cv.std())\n\nprint('XGBoost Tuning Validation')\nxgb_a_tuning_train = pd.DataFrame({\n    'model' : 'XGBoost tuning',\n    'RMSE_Score' : score_RMSE,\n    'Mean_RMSE': nilai_mean_RMSE,\n    'Std_RMSE': nilai_std_RMSE\n})\n\nxgb_a_tuning_train","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:08:35.289808Z","iopub.execute_input":"2022-08-08T13:08:35.290283Z","iopub.status.idle":"2022-08-08T13:08:43.557735Z","shell.execute_reply.started":"2022-08-08T13:08:35.290245Z","shell.execute_reply":"2022-08-08T13:08:43.557005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Evaluation XGBoost Tuning with X test","metadata":{}},{"cell_type":"code","source":"model_xgb_a_tuning = random_xgb.best_estimator_\n\nscore_rmse = []\nscore_r2 = [] \n\n# Prediction with X test\ny_pred_a_xgb = model_xgb_a_tuning.predict(X_test)\nscore_rmse.append(np.sqrt(mean_squared_error(y_test, y_pred_a_xgb)))\nscore_r2.append(r2_score(y_test, y_pred_a_xgb))\n\nxgb_a_tuning_test = pd.DataFrame({'model' : 'XGBoost After Tuning',\n                                    'RMSE': score_rmse, \n                                    'R-Squared' : score_r2\n                                   })\nxgb_a_tuning_test","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:08:43.559757Z","iopub.execute_input":"2022-08-08T13:08:43.561513Z","iopub.status.idle":"2022-08-08T13:08:43.593739Z","shell.execute_reply.started":"2022-08-08T13:08:43.561483Z","shell.execute_reply":"2022-08-08T13:08:43.592637Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Hyperparameter Tuning GradientBoost**","metadata":{}},{"cell_type":"code","source":"# The maximum depth limits the number of nodes in the tree. \nmax_depth = list(np.arange(15, 31))\n\n# Learning rate\nlearning_rate = list(np.arange(1, 100)/100)\n\n# The number of boosting stages to perform\nn_estimators = list(np.arange(100, 251))\n\n# The fraction of samples to be used for fitting the individual base learners.\nsubsample = list(np.arange(2, 10)/10)\n\n# The function to measure the quality of a split.\ncriterion = ['friedman_mse', 'squared_error', 'mse']\n\n# Hyperparam space Gradient Boost\nhyperparam_space_xgb = {\n    'model__max_depth': max_depth,\n    'model__learning_rate': learning_rate,\n    'model__n_estimators': n_estimators,\n    'model__subsample': subsample,\n    'model__criterion': criterion\n}","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:08:43.595811Z","iopub.execute_input":"2022-08-08T13:08:43.596824Z","iopub.status.idle":"2022-08-08T13:08:43.603412Z","shell.execute_reply.started":"2022-08-08T13:08:43.596794Z","shell.execute_reply":"2022-08-08T13:08:43.601976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Benchmark model with hyperparameter tuning\ngbr = GradientBoostingRegressor(random_state = 2022)\n\n# make algorithm chain\nestimator_gbr = Pipeline([\n        ('preprocessing', transformer),\n        ('model', gbr)\n        ])\n\ncrossval = KFold(n_splits=5, shuffle=True, random_state=1)\n\n# Hyperparameter tuning\nrandom_gbr = RandomizedSearchCV(\n    estimator_gbr, \n    param_distributions = hyperparam_space_xgb,\n    n_iter = 50,\n    cv = crossval, \n    scoring = ['neg_root_mean_squared_error', 'r2'], \n    n_jobs = -1,\n    refit = 'neg_root_mean_squared_error', # Just can optimize 1 metric\n    random_state = 2022 \n)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:08:43.604578Z","iopub.execute_input":"2022-08-08T13:08:43.604869Z","iopub.status.idle":"2022-08-08T13:08:43.617074Z","shell.execute_reply.started":"2022-08-08T13:08:43.604843Z","shell.execute_reply":"2022-08-08T13:08:43.616322Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"random_gbr.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:08:43.618160Z","iopub.execute_input":"2022-08-08T13:08:43.619099Z","iopub.status.idle":"2022-08-08T13:11:00.286914Z","shell.execute_reply.started":"2022-08-08T13:08:43.619056Z","shell.execute_reply":"2022-08-08T13:11:00.285864Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.DataFrame(random_gbr.cv_results_).sort_values(by = ['rank_test_neg_root_mean_squared_error']).head(5)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:11:00.288780Z","iopub.execute_input":"2022-08-08T13:11:00.289187Z","iopub.status.idle":"2022-08-08T13:11:00.322481Z","shell.execute_reply.started":"2022-08-08T13:11:00.289159Z","shell.execute_reply":"2022-08-08T13:11:00.321204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Gradient Boost')\nprint('Best_score:', random_gbr.best_score_*-1)\nprint('Best_params:', random_gbr.best_params_)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:11:00.324055Z","iopub.execute_input":"2022-08-08T13:11:00.324570Z","iopub.status.idle":"2022-08-08T13:11:00.335419Z","shell.execute_reply.started":"2022-08-08T13:11:00.324540Z","shell.execute_reply":"2022-08-08T13:11:00.334517Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_gbr_a_tuning = random_gbr.best_estimator_\ncrossval = KFold(n_splits=5, shuffle=True, random_state=2022)\n\nscore_RMSE = []\nnilai_mean_RMSE = []\nnilai_std_RMSE = []\n\nmodel_cv = (-cross_val_score(model_gbr_a_tuning, X_train, y_train, cv = crossval, scoring = 'neg_root_mean_squared_error'))\nscore_RMSE.append(model_cv)\nnilai_mean_RMSE.append(model_cv.mean())\nnilai_std_RMSE.append(model_cv.std())\n\nprint('Gradient Boost After Tuning Validation')\ngbr_a_tuning_train = pd.DataFrame({\n    'model' : 'Gradient Boost tuning',\n    'RMSE_Score' : score_RMSE,\n    'Mean_RMSE': nilai_mean_RMSE,\n    'Std_RMSE': nilai_std_RMSE\n})\ngbr_a_tuning_train","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:11:00.336464Z","iopub.execute_input":"2022-08-08T13:11:00.336761Z","iopub.status.idle":"2022-08-08T13:11:04.095260Z","shell.execute_reply.started":"2022-08-08T13:11:00.336735Z","shell.execute_reply":"2022-08-08T13:11:04.094326Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Evaluation Gradient Boost Tuning with X test","metadata":{}},{"cell_type":"code","source":"model_gbr_a_tuning = random_gbr.best_estimator_\n\nscore_rmse = []\nscore_r2 = [] \n\n# Prediction with X test\ny_pred_a_gbr = model_gbr_a_tuning.predict(X_test)\nscore_rmse.append(np.sqrt(mean_squared_error(y_test, y_pred_a_gbr)))\nscore_r2.append(r2_score(y_test, y_pred_a_gbr))\n\ngbr_a_tuning_test = pd.DataFrame({'model' : 'Gradient Boost After Tuning',\n                                    'RMSE': score_rmse, \n                                    'R-Squared' : score_r2\n                                   })\ngbr_a_tuning_test","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:11:04.096461Z","iopub.execute_input":"2022-08-08T13:11:04.097256Z","iopub.status.idle":"2022-08-08T13:11:04.124546Z","shell.execute_reply.started":"2022-08-08T13:11:04.097221Z","shell.execute_reply":"2022-08-08T13:11:04.123696Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Compare After Tuning and Before Tuning","metadata":{"tags":[],"toc-hr-collapsed":true}},{"cell_type":"markdown","source":"### Compare Evaluation Metric with cross val score (x train & y train)","metadata":{}},{"cell_type":"code","source":"pd.concat([best_b_tuning_train,\ngbr_a_tuning_train,\nxgb_a_tuning_train], axis = 0).sort_values('Mean_RMSE', ascending=True).reset_index().drop('index', axis = 1 )","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:11:04.126934Z","iopub.execute_input":"2022-08-08T13:11:04.128054Z","iopub.status.idle":"2022-08-08T13:11:04.143564Z","shell.execute_reply.started":"2022-08-08T13:11:04.128015Z","shell.execute_reply":"2022-08-08T13:11:04.142727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Compare Evaluation Metric (x test)","metadata":{}},{"cell_type":"code","source":"pd.concat([best_b_tuning_test,\ngbr_a_tuning_test,\nxgb_a_tuning_test], axis = 0).sort_values('RMSE', ascending=True).reset_index().drop('index', axis = 1 )","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:11:04.147176Z","iopub.execute_input":"2022-08-08T13:11:04.148502Z","iopub.status.idle":"2022-08-08T13:11:04.165202Z","shell.execute_reply.started":"2022-08-08T13:11:04.148453Z","shell.execute_reply":"2022-08-08T13:11:04.163967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Gradient Boost Before Tuning** has bigger score RMSE and R2 , but **Gradient Boost After Tuning** has more stable score RMSE, We Choose to compare **Gradient Boost Before Tuning** and **Gradient Boost After Tuning** with plot","metadata":{}},{"cell_type":"markdown","source":"<!-- We choose **Gradient Boost After Tuning** as the best model, even Gradient Boost before tunning have bigger score in RMSE Xtest (lower is better), but in cros val score **Gradient Boost After Tuning have a stable rmse score** -->","metadata":{}},{"cell_type":"markdown","source":"#  Compare Model Predict with Actual\n---","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(26, 8))\nplt.subplot(1,2,1)\nplot = sns.regplot(x=y_test, y=y_pred_b_GradientBoost, \n                   scatter_kws = {'color' : 'g'},\n                   line_kws = {'color' : 'r'}).set(title='Plot A : Gradient Boost Before Tuning\\nActual vs. Prediction Price', \n                                               xlabel='Actual Price', \n                                               ylabel='Predicted Price');\nplt.subplot(1,2,2)\nplot = sns.regplot(x=y_test, y=y_pred_a_gbr, \n                   scatter_kws = {'color' : 'g'},\n                   line_kws = {'color' : 'r'}).set(title='Plot B : Gradient Boost After Tuning\\nActual vs. Prediction Price', \n                                               xlabel='Actual Price', \n                                               ylabel='Predicted Price');","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:11:04.166950Z","iopub.execute_input":"2022-08-08T13:11:04.167480Z","iopub.status.idle":"2022-08-08T13:11:04.736139Z","shell.execute_reply.started":"2022-08-08T13:11:04.167451Z","shell.execute_reply":"2022-08-08T13:11:04.735033Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"glance plot a and plot b plots look the same, but if we see detail in right corner there has 1 data, at plot a distance actual and predict look lower than plot b","metadata":{}},{"cell_type":"markdown","source":"**We choose Gradient Boost be best model in this case**","metadata":{}},{"cell_type":"markdown","source":"## Compare Gradient Boost Before Tuning predict with actual","metadata":{"tags":[]}},{"cell_type":"code","source":"test = pd.DataFrame({'Predicted':y_pred_b_GradientBoost,'Actual':y_test})\nfig= plt.figure(figsize=(16,8))\ntest = test.reset_index()\ntest = test.drop(['index'],axis=1)\nplt.plot(test[:150])\nplt.title('Compare  predict dengan Actual', size = 16)\nplt.legend(['Actual','Predicted'])\nsns.jointplot(x='Actual',y='Predicted',data=test,kind=\"reg\")","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:11:04.739071Z","iopub.execute_input":"2022-08-08T13:11:04.739467Z","iopub.status.idle":"2022-08-08T13:11:05.481531Z","shell.execute_reply.started":"2022-08-08T13:11:04.739441Z","shell.execute_reply":"2022-08-08T13:11:05.480278Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Incorrect Predict Gradient Boost Before Tuning","metadata":{}},{"cell_type":"code","source":"test['Different'] = test['Actual'] - test['Predicted']","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:11:05.483611Z","iopub.execute_input":"2022-08-08T13:11:05.484282Z","iopub.status.idle":"2022-08-08T13:11:05.489489Z","shell.execute_reply.started":"2022-08-08T13:11:05.484221Z","shell.execute_reply":"2022-08-08T13:11:05.488461Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test['Different'] = test['Different'].abs()","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:11:05.490932Z","iopub.execute_input":"2022-08-08T13:11:05.491185Z","iopub.status.idle":"2022-08-08T13:11:05.504709Z","shell.execute_reply.started":"2022-08-08T13:11:05.491160Z","shell.execute_reply":"2022-08-08T13:11:05.503949Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test[test['Different'] == test['Different'].max()]","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:11:05.506512Z","iopub.execute_input":"2022-08-08T13:11:05.507640Z","iopub.status.idle":"2022-08-08T13:11:05.527207Z","shell.execute_reply.started":"2022-08-08T13:11:05.507609Z","shell.execute_reply":"2022-08-08T13:11:05.525933Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.iloc[139]","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:11:05.530037Z","iopub.execute_input":"2022-08-08T13:11:05.532614Z","iopub.status.idle":"2022-08-08T13:11:05.544567Z","shell.execute_reply.started":"2022-08-08T13:11:05.532569Z","shell.execute_reply":"2022-08-08T13:11:05.543004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Test Machine Learning","metadata":{}},{"cell_type":"code","source":"import pickle","metadata":{"tags":[],"execution":{"iopub.status.busy":"2022-08-08T13:11:05.548072Z","iopub.execute_input":"2022-08-08T13:11:05.548412Z","iopub.status.idle":"2022-08-08T13:11:05.555571Z","shell.execute_reply.started":"2022-08-08T13:11:05.548384Z","shell.execute_reply":"2022-08-08T13:11:05.554650Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Approuching Data Test like Data Train","metadata":{}},{"cell_type":"code","source":"df_test","metadata":{"tags":[],"execution":{"iopub.status.busy":"2022-08-08T13:11:05.557837Z","iopub.execute_input":"2022-08-08T13:11:05.558191Z","iopub.status.idle":"2022-08-08T13:11:05.597488Z","shell.execute_reply.started":"2022-08-08T13:11:05.558167Z","shell.execute_reply":"2022-08-08T13:11:05.596511Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Handling Missing Value","metadata":{}},{"cell_type":"code","source":"# check missing value\nmiss = pd.DataFrame({'Missing Value' : df_test.isna().sum()})\nmiss[miss['Missing Value'] > 0 ]","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:11:05.598483Z","iopub.execute_input":"2022-08-08T13:11:05.598785Z","iopub.status.idle":"2022-08-08T13:11:05.615536Z","shell.execute_reply.started":"2022-08-08T13:11:05.598752Z","shell.execute_reply":"2022-08-08T13:11:05.614715Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"s_imputer_1 = SimpleImputer(strategy='constant', fill_value=0)\ns_imputer_2 = SimpleImputer(strategy='constant', fill_value=\"None\")\ns_imputer_3 = SimpleImputer(strategy='most_frequent')\ns_imputer_4 = SimpleImputer(strategy='constant', fill_value=\"NA\")\ns_imputer_5 = SimpleImputer(strategy='constant', fill_value=\"Other\")","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:11:05.627486Z","iopub.execute_input":"2022-08-08T13:11:05.628795Z","iopub.status.idle":"2022-08-08T13:11:05.636644Z","shell.execute_reply.started":"2022-08-08T13:11:05.628732Z","shell.execute_reply":"2022-08-08T13:11:05.635473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# filling missing value with iterative imputer\nfeat_1 = ['LotFrontage', 'MasVnrArea', 'GarageYrBlt', 'BsmtFinSF1', 'BsmtFinSF2', 'BsmtUnfSF', 'TotalBsmtSF', 'BsmtFullBath', 'BsmtHalfBath', 'KitchenQual', 'GarageCars', 'GarageArea']\nfeat_2 = ['MasVnrType']\nfeat_3 = ['Electrical', 'MSZoning', 'Utilities', 'Functional', 'SaleType']\nfeat_4 = ['Alley', 'BsmtQual', 'BsmtCond', 'BsmtExposure', 'BsmtFinType1', \n          'BsmtFinType2', 'FireplaceQu', 'GarageType', 'GarageFinish', \n          'GarageQual', 'GarageCond', 'PoolQC', 'Fence', 'MiscFeature']\nfeat_5 = ['Exterior1st', 'Exterior2nd']","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:11:05.638444Z","iopub.execute_input":"2022-08-08T13:11:05.638888Z","iopub.status.idle":"2022-08-08T13:11:05.652022Z","shell.execute_reply.started":"2022-08-08T13:11:05.638861Z","shell.execute_reply":"2022-08-08T13:11:05.651056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test[feat_1] = s_imputer_1.fit_transform(df_test[feat_1])\ndf_test[feat_2] = s_imputer_2.fit_transform(df_test[feat_2])\ndf_test[feat_3] = s_imputer_3.fit_transform(df_test[feat_3])\ndf_test[feat_4] = s_imputer_4.fit_transform(df_test[feat_4])\ndf_test[feat_5] = s_imputer_5.fit_transform(df_test[feat_5])","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:11:05.654298Z","iopub.execute_input":"2022-08-08T13:11:05.655692Z","iopub.status.idle":"2022-08-08T13:11:05.699297Z","shell.execute_reply.started":"2022-08-08T13:11:05.655630Z","shell.execute_reply":"2022-08-08T13:11:05.697775Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check missing value\nmiss = pd.DataFrame({'Missing Value' : df_test.isna().sum()})\nmiss[miss['Missing Value'] > 0 ]","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:11:05.700835Z","iopub.execute_input":"2022-08-08T13:11:05.701285Z","iopub.status.idle":"2022-08-08T13:11:05.716829Z","shell.execute_reply.started":"2022-08-08T13:11:05.701249Z","shell.execute_reply":"2022-08-08T13:11:05.715722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Missing value heatmap\nplt.figure(figsize = (15, 10))\nsns.heatmap(df_test.isnull(), cbar=False);\nplt.title('Heatmap Missing Value \\n', size =20)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:11:05.718559Z","iopub.execute_input":"2022-08-08T13:11:05.719227Z","iopub.status.idle":"2022-08-08T13:11:07.117180Z","shell.execute_reply.started":"2022-08-08T13:11:05.719199Z","shell.execute_reply":"2022-08-08T13:11:07.116307Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# index_missing = list(df_test[df_test.isna().any(axis=1)].index)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:11:07.118445Z","iopub.execute_input":"2022-08-08T13:11:07.118957Z","iopub.status.idle":"2022-08-08T13:11:07.123432Z","shell.execute_reply.started":"2022-08-08T13:11:07.118922Z","shell.execute_reply":"2022-08-08T13:11:07.122283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# index_missing = []\n# missing_feat_test = ['MSZoning', 'Utilities', 'Exterior1st', 'Exterior2nd', 'BsmtFinSF1', \n#                 'BsmtFinSF2', 'BsmtUnfSF', 'TotalBsmtSF', 'BsmtFullBath', 'BsmtHalfBath', \n#                 'KitchenQual', 'Functional', 'GarageCars', 'GarageArea', 'SaleType']\n# for i in missing_feat_test :\n#     index_missing.extend(list(df_test[df_test[i].isna() == True].index))\n\n# index_missing = tuple(index_missing)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:11:07.125256Z","iopub.execute_input":"2022-08-08T13:11:07.125737Z","iopub.status.idle":"2022-08-08T13:11:07.138330Z","shell.execute_reply.started":"2022-08-08T13:11:07.125700Z","shell.execute_reply":"2022-08-08T13:11:07.137537Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# index_missing","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:11:07.139429Z","iopub.execute_input":"2022-08-08T13:11:07.139815Z","iopub.status.idle":"2022-08-08T13:11:07.154956Z","shell.execute_reply.started":"2022-08-08T13:11:07.139790Z","shell.execute_reply":"2022-08-08T13:11:07.153746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# df_test.iloc[[95, 455, 485, 660, 691, 728, 756, 790, 1013, 1029, 1116, 1444]]","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:11:07.156603Z","iopub.execute_input":"2022-08-08T13:11:07.156948Z","iopub.status.idle":"2022-08-08T13:11:07.169862Z","shell.execute_reply.started":"2022-08-08T13:11:07.156923Z","shell.execute_reply":"2022-08-08T13:11:07.168471Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# df_test.dropna(inplace = True)","metadata":{"tags":[],"execution":{"iopub.status.busy":"2022-08-08T13:11:07.171407Z","iopub.execute_input":"2022-08-08T13:11:07.171944Z","iopub.status.idle":"2022-08-08T13:11:07.184605Z","shell.execute_reply.started":"2022-08-08T13:11:07.171909Z","shell.execute_reply":"2022-08-08T13:11:07.183440Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<!-- we drop row the remaining missing Value (12 data) -->","metadata":{}},{"cell_type":"code","source":"# # check missing value\n# miss = pd.DataFrame({'Missing Value' : df_test.isna().sum()})\n# miss[miss['Missing Value'] > 0 ]","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:11:07.185820Z","iopub.execute_input":"2022-08-08T13:11:07.186182Z","iopub.status.idle":"2022-08-08T13:11:07.201786Z","shell.execute_reply.started":"2022-08-08T13:11:07.186157Z","shell.execute_reply":"2022-08-08T13:11:07.200509Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Approaching Categorical Features**","metadata":{}},{"cell_type":"code","source":"for i in object_feat:\n    df_test[i] = le.fit_transform(df_test[i].astype(str))\n\nprint (df_test.info())","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:11:07.203171Z","iopub.execute_input":"2022-08-08T13:11:07.203894Z","iopub.status.idle":"2022-08-08T13:11:07.273798Z","shell.execute_reply.started":"2022-08-08T13:11:07.203866Z","shell.execute_reply":"2022-08-08T13:11:07.273113Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Save ML","metadata":{}},{"cell_type":"code","source":"file_name = 'Price Predict.sav'\ngbr = GradientBoostingRegressor(random_state = 2022)\nmodel = Pipeline([\n    ('preprocessing', transformer),\n    ('model', gbr)\n    ])\nmodel.fit(X_train, y_train)\npickle.dump(model, open(file_name,'wb'))","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:11:07.274972Z","iopub.execute_input":"2022-08-08T13:11:07.275231Z","iopub.status.idle":"2022-08-08T13:11:07.838980Z","shell.execute_reply.started":"2022-08-08T13:11:07.275205Z","shell.execute_reply":"2022-08-08T13:11:07.837971Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Load ML","metadata":{}},{"cell_type":"code","source":"loaded_model = pickle.load(open(file_name,'rb'))","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:11:07.840257Z","iopub.execute_input":"2022-08-08T13:11:07.840516Z","iopub.status.idle":"2022-08-08T13:11:07.849252Z","shell.execute_reply.started":"2022-08-08T13:11:07.840490Z","shell.execute_reply":"2022-08-08T13:11:07.848160Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predict = loaded_model.predict(df_test)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:11:07.850818Z","iopub.execute_input":"2022-08-08T13:11:07.851915Z","iopub.status.idle":"2022-08-08T13:11:07.879615Z","shell.execute_reply.started":"2022-08-08T13:11:07.851876Z","shell.execute_reply":"2022-08-08T13:11:07.878250Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_submission = pd.DataFrame({'Id' : df_test['Id'], 'SalePrice' : predict})\nsample_submission","metadata":{"execution":{"iopub.status.busy":"2022-08-08T13:11:07.881472Z","iopub.execute_input":"2022-08-08T13:11:07.881923Z","iopub.status.idle":"2022-08-08T13:11:07.893723Z","shell.execute_reply.started":"2022-08-08T13:11:07.881884Z","shell.execute_reply":"2022-08-08T13:11:07.892744Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Conclusion","metadata":{}},{"cell_type":"markdown","source":"RMSE results show machine learning will 24216.4. we can conclude that when this model is used to predict house prices in Ames in the value range as trained on the model, the average estimate will miss about 24,216.4. but in fact this model has biggest miss 153,222. and R2 score 0.9, which means that the independent variable greatly affects the dependent variable","metadata":{}}]}