{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.7.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":5407,"databundleVersionId":868283,"sourceType":"competition"}],"dockerImageVersionId":30458,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Final Project\n\nAs in all machine learning problems you should complete the following steps:\n1. Load and explore the data (using plots and histograms )\n2. Clean/preprocess/transform the data if necessary (first performed on training data and next on test data)\n3. Train the machine learning model (in this assignment two models, one linear and one non-linear)\n4. Evaluate and optimise the model\n\nHowever, while performing the steps you should consider our main goals and research questions for doing this project and try to adress them. This can be either in each step or afterwards in discussion and conclusion section (your call!). Below are the research question we are interested in:\n\n* What are the necessary step to clean and prepare the dataset as we have categorial features and missing data.\n* What model/classifier provides the best result for this application.\n* What is the best choice for our cost function and performance metrics for this problem.\n\nPlease also note the following points during the assignment:\n\n* Use functions from open source libraries like sci-kit learn and keras and avoid using your own hand-written functions from previous assignments. You can also use our [cheatsheet](https://colab.research.google.com/drive/12h-QBlsaWXkjGIRoJXfF4yqi1elnX9qn?usp=sharing).\n* Feel free to contact us on Teams if you need more description.\n* This notebook is structured like a scientific paper. The text should provide a high-level overview of your approach. Please don't include any details about your code in the text but add them as comments in the code itself. Your code should be cleane and readable with enough comments.\n* There are some instructions and questions in each section, remove the highlighted text in blue and replace it with your explanations and answers.\n\nYou can delete this section before submission.","metadata":{"id":"GF4WsCxg-lL9"}},{"cell_type":"markdown","source":"## 1. Introduction\n","metadata":{"id":"3NnAgB-y0knj"}},{"cell_type":"markdown","source":"Rutger Gosselink 5770289 rutgergosselink\nArd Geuze 5743052 ardgeuze\n","metadata":{"id":"DPJV5e8sDVnT"}},{"cell_type":"markdown","source":"## 2. Data\n","metadata":{"id":"PB7FLdQn-dmK"}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd \nfrom sklearn.model_selection import train_test_split\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nfrom sklearn.linear_model import LinearRegression\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"execution":{"iopub.status.busy":"2024-04-29T12:36:45.804746Z","iopub.execute_input":"2024-04-29T12:36:45.805137Z","iopub.status.idle":"2024-04-29T12:36:45.817499Z","shell.execute_reply.started":"2024-04-29T12:36:45.805103Z","shell.execute_reply":"2024-04-29T12:36:45.816384Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2.1 Dataset\n<font color=#6698FF>In this section, we load and explore the dataset.","metadata":{"id":"vfGikOxUAxJB"}},{"cell_type":"code","source":"train_file_path = \"../input/house-prices-advanced-regression-techniques/train.csv\"\ndataset_df = pd.read_csv(train_file_path)\nprint(\"Full train dataset shape is {}\".format(dataset_df.shape))\n\n\nprint(dataset_df['SalePrice'].describe())\n\n\n\n","metadata":{"execution":{"iopub.status.busy":"2024-04-29T12:36:45.819798Z","iopub.execute_input":"2024-04-29T12:36:45.820735Z","iopub.status.idle":"2024-04-29T12:36:45.860073Z","shell.execute_reply.started":"2024-04-29T12:36:45.820691Z","shell.execute_reply":"2024-04-29T12:36:45.858972Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"###### 2.1.1 Train-test split\n<font color=#6698FF>In the below, we split the train data into a test and a train set. Set a value for the `test_size` yourself. Argue why the test value can not be too small or too large. You can also use k-fold cross validation.\n","metadata":{"id":"udqAFg1C3JHe"}},{"cell_type":"markdown","source":"The test size needs to be big enough that the model can be tested without the variance of the test data being a problem. If the test size is to large, there is not enough data left to train the model. ","metadata":{}},{"cell_type":"code","source":"df_num = np.array(dataset_df)\nprint(np.shape(df_num))\nX = df_num[:,:36]\ny = df_num[:,36]\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2,shuffle=True, random_state=42)\nprint(np.shape(X_train))\nprint(type(X_train))","metadata":{"execution":{"iopub.status.busy":"2024-04-29T12:36:45.861467Z","iopub.execute_input":"2024-04-29T12:36:45.861769Z","iopub.status.idle":"2024-04-29T12:36:45.875535Z","shell.execute_reply.started":"2024-04-29T12:36:45.861739Z","shell.execute_reply":"2024-04-29T12:36:45.874418Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2.2 Data Exploration\n\n<font color=#6698FF> Explore the features and target variables of the dataset. Think about making some scatter plots, box plots, histograms or printing the data, but feel free to choose any method that suits you.\nWhat do you think is the right performance\nmetric to use for this dataset? Clearly explain which performance metric you\nchoose and why.\nAlgorithmic bias can be a real problem in Machine Learning.  Explain what you believe.</font>","metadata":{"id":"QN5jS2_p_Y9N"}},{"cell_type":"markdown","source":"The best performance metric is the error with a logaritm. This will compenste for the big los when the estimated value is very large, but it is a few percent of.","metadata":{}},{"cell_type":"code","source":"\ntrain_file_path = \"../input/house-prices-advanced-regression-techniques/train.csv\"\ndataset_df = pd.read_csv(train_file_path)\nprint(\"Full train dataset shape is {}\".format(dataset_df.shape))\n\n\nprint(dataset_df['SalePrice'].describe())\nplt.figure(figsize=(9, 8))\nsns.distplot(dataset_df['SalePrice'], color='g', bins=100, hist_kws={'alpha': 0.4});\n\ndf_num = dataset_df.select_dtypes(include = ['float64', 'int64'])\n\ndf_num.hist(figsize=(16, 20), bins=50, xlabelsize=8, ylabelsize=8);\n","metadata":{"execution":{"iopub.status.busy":"2024-04-29T12:36:45.876817Z","iopub.execute_input":"2024-04-29T12:36:45.877276Z","iopub.status.idle":"2024-04-29T12:36:54.939222Z","shell.execute_reply.started":"2024-04-29T12:36:45.877237Z","shell.execute_reply":"2024-04-29T12:36:54.938200Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2.3 Data Preparation\n\n<font color=#6698FF>  This dataset hasn’t been cleaned yet. Meaning that some attributes (features) are in numerical format and some are in categorial format. Moreover, there are missing values as well. However, all Scikit-learn’s implementations of these algorithms expect numerical features. Check for all features if they are in categorial and use a method to transform them to numerical values. For the numerical data, handle the missing data and normalize the data. \nNote that you are only allowed to use training data for preprocessing but you then need to perform similar changes on test data too.\nYou can use [pipelining](https://scikit-learn.org/stable/modules/generated/sklearn.pipeline.Pipeline.html) to help with the preprocessing.</font>","metadata":{"id":"4A5GsVzgATRu"}},{"cell_type":"code","source":"dataset_df = dataset_df.drop('Id', axis=1)\ndf_num = dataset_df.select_dtypes(include = ['float64', 'int64'])\ndf_num = df_num.fillna(0)\ndf_num = df_num.fillna(0)\ndf_num = np.array(df_num)\nprint(np.shape(df_num))\nX = df_num[:,:36]\ny = df_num[:,36]\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2,shuffle=True, random_state=42)\nprint(np.shape(X_train))\nprint(type(X_train))","metadata":{"execution":{"iopub.status.busy":"2024-04-29T12:36:54.941875Z","iopub.execute_input":"2024-04-29T12:36:54.942323Z","iopub.status.idle":"2024-04-29T12:36:54.959563Z","shell.execute_reply.started":"2024-04-29T12:36:54.942285Z","shell.execute_reply":"2024-04-29T12:36:54.958434Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n## 3. Training and Results\n<font color=#6698FF> Briefly introduce algorithms you choose. \nPresent your balanced accuracies for both test and training data for all classifiers. Analyse the performance on test and training in terms of bias and variance. Give one advantage and one drawback of the method you use.\n","metadata":{"id":"IUcc2ue7yrOv"}},{"cell_type":"markdown","source":"PCA can improve the time to train a model, but can ruduce accuracy","metadata":{}},{"cell_type":"code","source":"# linear\nplt.figure(1)\nfrom sklearn.decomposition import PCA\nscore =[]\nfor i in range(36):\n    pca = PCA(n_components=i+1)\n    pca = pca.fit(X_train)\n    X_train_pca=pca.transform(X_train);\n    X_test_pca=pca.transform(X_test)\n\n\n    model = LinearRegression()\n    model.fit(X_train_pca, y_train)\n    score.append(model.score(X_test_pca, y_test))\n\n    print(model.score(X_test_pca, y_test))\nplt.plot(score)\nplt.show\n\n# non linear\nplt.figure(2)\nfrom sklearn.ensemble import RandomForestClassifier\n\n\nfrom sklearn import preprocessing\nscaler = preprocessing.StandardScaler().fit(X_train)\nX_trains = scaler.transform(X_train)\nX_tests = scaler.transform(X_test)\n\nmodel = RandomForestClassifier(1000, max_depth = 3, max_features = 3)\n\nmodel.fit(X_trains, y_train)\n\nprint(model.score(X_tests, y_test))\nimportance = model.feature_importances_\n\n#creat a dictionary with key=indices, and values=importance\nimportant_features_dict = {}\nfor idx, val in enumerate(importance):\n    important_features_dict[idx] = val\n# sorting\nimportant_features_list = sorted(important_features_dict,\n                                 key=important_features_dict.get,\n                                 reverse=True)\nimportant_features_list[:5]\nplt.bar([x for x in range(len(importance))], importance)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-29T12:36:54.960900Z","iopub.execute_input":"2024-04-29T12:36:54.961254Z","iopub.status.idle":"2024-04-29T12:37:00.567127Z","shell.execute_reply.started":"2024-04-29T12:36:54.961223Z","shell.execute_reply":"2024-04-29T12:37:00.566033Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4. Discussion and Conclusion\n\n<font color=#6698FF>Discuss all the choices you made during the process and your final conclusions.  Highlight the strong points of your approach, discuss its shortcomings  and suggest some future approaches that may improve it. Please be self critical here. The assignment is not about achieving a state of the art performance, but about showing what you have learned the concepts during the course.</font>","metadata":{"id":"yekgnxu7y0gE"}},{"cell_type":"markdown","source":"First all non numerical values are removed from the data set. There is one numerical varable called id that has no meaning when trained, so this is also removed. Then the X and Y values are seperated from eachother and then seperated into train and test data. \n\nFor the linear model PCA is used beforehand. The effect of the variable count for this PCA is shown in the graph. after 25 it does not improve much.\n\nfor the random forest clasifier the results are garbage. the best result came with a high number of trees. 1000 was chosen for this. A max depth of 3 and 3 max features gave the best results without overfitting.","metadata":{}},{"cell_type":"markdown","source":"## 5. References","metadata":{"id":"wn7PBEmWESIi"}},{"cell_type":"markdown","source":"https://www.kaggle.com/code/prathamaggarwal20/house-price-prediction\nhttps://www.kaggle.com/code/gusthema/house-prices-prediction-using-tfdf/notebook\nhttps://colab.research.google.com/drive/12h-QBlsaWXkjGIRoJXfF4yqi1elnX9qn?usp=sharing#scrollTo=91o60OSRsxmf\nhttps://scikit-learn.org/stable/modules/generated/sklearn.ensemble.RandomForestClassifier.html\nhttps://www.quora.com/Why-is-my-Random-Forest-implementation-showing-poor-and-kind-of-weird-performance-on-training-data-and-much-worse-on-testing-data","metadata":{}}]}