{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":50160,"databundleVersionId":7921029,"sourceType":"competition"}],"dockerImageVersionId":30698,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# COS20083 Advanced Data Analytics\n\n## Assignment 2: Case Study and Algorithm Implementation\n\n### Semester 1, 2024","metadata":{}},{"cell_type":"markdown","source":"#### Group Number: <p style =\"color: red;\">13</p>\n#### Group Members: \n<p style =\"color: red;\">101234017 Ye Ding NG</p>\n<p style =\"color: red;\">102767857 Nicholas Tiew Leong TING</p>\n","metadata":{}},{"cell_type":"markdown","source":"## <p style =\"color: blue;\">1. Introduction</p>","metadata":{}},{"cell_type":"markdown","source":"The purpose of this assignment is to give a simulated real-world data science scenario to students, to learn the full steps of a data science project, and to learn how to collaboration and manage tasks effectively as a team.\n\nThe project is based on an online on-going competition hosted by Kaggle. The \"Home Credit - Credit Risk Model Stability\" competition on Kaggle aims to improve the accuracy and stability of credit risk models. Home Credit provides loans to clients with limited credit history, requiring robust predictive models to assess risk accurately. Participants are tasked with developing models that can maintain high performance despite potential changes in economic conditions and customer behavior. The goal is to ensure the models are both effective and resilient, enhancing the company's ability to make informed lending decisions.\n\nIn summary, each team's overall project-process will be to understand the project set by Kaggle, perform data exploration, data analysis, model building and finally to test their model and evaluate the performance.","metadata":{}},{"cell_type":"markdown","source":"## <p style =\"color: blue;\">2. Data Collection</p>","metadata":{}},{"cell_type":"markdown","source":"Before performing any form of data analysis and data processing, the fundamental goal of data collection is to take our time and fully understand the data that we are working with. The team will first import all the datasets by merging, followed by understanding how and when it will be used.","metadata":{}},{"cell_type":"code","source":"# Reference: https://www.kaggle.com/code/aaachen/home-credit-clean-code-lightgbm/notebook\nimport warnings\nwarnings.simplefilter(action='ignore', category=FutureWarning)\n\n\nimport os\nimport gc\nimport numpy as np\nimport pandas as pd\nimport polars as pl\nprint(pl.__version__)\nfrom glob import glob\nfrom pathlib import Path\nfrom datetime import datetime\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom sklearn.model_selection import TimeSeriesSplit, GroupKFold, StratifiedGroupKFold\nfrom sklearn.base import BaseEstimator, RegressorMixin\nfrom sklearn.metrics import mean_squared_error\nfrom sklearn import metrics\nfrom sklearn.model_selection import train_test_split, LeaveOneOut, KFold, cross_val_score\nfrom sklearn.preprocessing import PolynomialFeatures\nfrom sklearn.metrics import roc_auc_score,roc_curve\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.metrics import classification_report\nfrom sklearn.metrics import mean_squared_error\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.neighbors import KNeighborsRegressor\nfrom sklearn.tree import DecisionTreeRegressor\nfrom xgboost import XGBRegressor\nimport lightgbm as lgb\n\ndataPath = \"home-credit-credit-risk-model-stability/\"","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:17:16.583209Z","iopub.execute_input":"2024-05-31T06:17:16.583578Z","iopub.status.idle":"2024-05-31T06:17:19.252544Z","shell.execute_reply.started":"2024-05-31T06:17:16.583538Z","shell.execute_reply":"2024-05-31T06:17:19.251651Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class Pipeline:\n    @staticmethod\n    def set_table_dtypes(df):\n        for col in df.columns:\n            if col in [\"case_id\", \"WEEK_NUM\", \"num_group1\", \"num_group2\"]:\n                df = df.with_columns(pl.col(col).cast(pl.Int64))\n            elif col in [\"date_decision\"]:\n                df = df.with_columns(pl.col(col).cast(pl.Date))\n            elif col[-1] in (\"P\", \"A\"):\n                df = df.with_columns(pl.col(col).cast(pl.Float64))\n            elif col[-1] in (\"M\",):\n                df = df.with_columns(pl.col(col).cast(pl.String))\n            elif col[-1] in (\"D\",):\n                df = df.with_columns(pl.col(col).cast(pl.Date))\n        return df\n\n    @staticmethod\n    def handle_dates(df):\n        for col in df.columns:\n            if col[-1] in (\"D\",):\n                df = df.with_columns(pl.col(col) - pl.col(\"date_decision\"))\n                df = df.with_columns(pl.col(col).dt.total_days())\n        df = df.drop(\"date_decision\", \"MONTH\")\n\n        return df\n\n    @staticmethod\n    def filter_cols(df):\n        for col in df.columns:\n            if col not in [\"target\", \"case_id\", \"WEEK_NUM\"]:\n                isnull = df[col].is_null().mean()\n\n                if isnull > 0.95:\n                    df = df.drop(col)\n\n        for col in df.columns:\n            if (col not in [\"target\", \"case_id\", \"WEEK_NUM\"]) & (df[col].dtype == pl.String):\n                freq = df[col].n_unique()\n\n                if (freq == 1) | (freq > 200):\n                    df = df.drop(col)\n\n        return df","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:17:19.254233Z","iopub.execute_input":"2024-05-31T06:17:19.254642Z","iopub.status.idle":"2024-05-31T06:17:19.269135Z","shell.execute_reply.started":"2024-05-31T06:17:19.254604Z","shell.execute_reply":"2024-05-31T06:17:19.268461Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class Aggregator:\n    @staticmethod\n    def num_expr(df):\n        cols = [col for col in df.columns if col[-1] in (\"P\", \"A\")]\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]\n        return expr_max\n\n    @staticmethod\n    def date_expr(df):\n        cols = [col for col in df.columns if col[-1] in (\"D\",)]\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]\n        return expr_max\n\n    @staticmethod\n    def str_expr(df):\n        cols = [col for col in df.columns if col[-1] in (\"M\",)]\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]\n        return expr_max\n\n    @staticmethod\n    def other_expr(df):\n        cols = [col for col in df.columns if col[-1] in (\"T\", \"L\")]\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]\n        return expr_max\n\n    @staticmethod\n    def count_expr(df):\n        cols = [col for col in df.columns if \"num_group\" in col]\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]\n        return expr_max\n\n    @staticmethod\n    def get_exprs(df):\n        exprs = Aggregator.num_expr(df) + \\\n                Aggregator.date_expr(df) + \\\n                Aggregator.str_expr(df) + \\\n                Aggregator.other_expr(df) + \\\n                Aggregator.count_expr(df)\n        return exprs","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:17:19.270442Z","iopub.execute_input":"2024-05-31T06:17:19.270720Z","iopub.status.idle":"2024-05-31T06:17:19.281931Z","shell.execute_reply.started":"2024-05-31T06:17:19.270698Z","shell.execute_reply":"2024-05-31T06:17:19.280846Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Read a file\ndef read_file(path, depth=None):\n    df = pl.read_parquet(path)\n    df = df.pipe(Pipeline.set_table_dtypes)\n    if depth in [1, 2]:\n        df = df.group_by(\"case_id\").agg(Aggregator.get_exprs(df))\n    return df\n\n# Read multiple files\ndef read_files(regex_path, depth=None):\n    chunks = []\n    for path in glob(str(regex_path)):\n        chunks.append(pl.read_parquet(path).pipe(Pipeline.set_table_dtypes))\n    df = pl.concat(chunks, how=\"vertical_relaxed\")\n    if depth in [1, 2]:\n        df = df.group_by(\"case_id\").agg(Aggregator.get_exprs(df))\n    return df","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:17:19.283030Z","iopub.execute_input":"2024-05-31T06:17:19.283318Z","iopub.status.idle":"2024-05-31T06:17:19.295542Z","shell.execute_reply.started":"2024-05-31T06:17:19.283294Z","shell.execute_reply":"2024-05-31T06:17:19.294859Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Merge the dataset\ndef feature_eng(df_base, depth_0, depth_1, depth_2):\n    df_base = (\n        df_base\n        .with_columns(\n            month_decision = pl.col(\"date_decision\").dt.month(),\n            weekday_decision = pl.col(\"date_decision\").dt.weekday(),\n        )\n    )\n    for i, df in enumerate(depth_0 + depth_1 + depth_2):\n        df_base = df_base.join(df, how=\"left\", on=\"case_id\", suffix=f\"_{i}\")\n    df_base = df_base.pipe(Pipeline.handle_dates)\n    return df_base","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:17:19.297791Z","iopub.execute_input":"2024-05-31T06:17:19.298564Z","iopub.status.idle":"2024-05-31T06:17:19.306724Z","shell.execute_reply.started":"2024-05-31T06:17:19.298541Z","shell.execute_reply":"2024-05-31T06:17:19.306039Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Convert polar object to panda object\ndef to_pandas(df_data, cat_cols=None):\n    df_data = df_data.to_pandas()\n    if cat_cols is None:\n        cat_cols = list(df_data.select_dtypes(\"object\").columns)\n    df_data[cat_cols] = df_data[cat_cols].astype(\"category\")\n    return df_data, cat_cols","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:17:19.307816Z","iopub.execute_input":"2024-05-31T06:17:19.308301Z","iopub.status.idle":"2024-05-31T06:17:19.318171Z","shell.execute_reply.started":"2024-05-31T06:17:19.308278Z","shell.execute_reply":"2024-05-31T06:17:19.317556Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ROOT            = Path(\"/kaggle/input/home-credit-credit-risk-model-stability\")\n# ROOT            = Path(\"./input\")\n\nTRAIN_DIR       = ROOT / \"parquet_files\" / \"train\"\nTEST_DIR        = ROOT / \"parquet_files\" / \"test\"","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:17:19.319316Z","iopub.execute_input":"2024-05-31T06:17:19.319816Z","iopub.status.idle":"2024-05-31T06:17:19.327132Z","shell.execute_reply.started":"2024-05-31T06:17:19.319793Z","shell.execute_reply":"2024-05-31T06:17:19.326476Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Read the training dataset file\ndata_store = {\n    \"df_base\": read_file(TRAIN_DIR / \"train_base.parquet\"),\n    \"depth_0\": [\n        read_file(TRAIN_DIR / \"train_static_cb_0.parquet\"),\n        read_files(TRAIN_DIR / \"train_static_0_*.parquet\"),\n    ],\n    \"depth_1\": [\n        read_files(TRAIN_DIR / \"train_applprev_1_*.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_tax_registry_a_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_tax_registry_b_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_tax_registry_c_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_credit_bureau_b_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_other_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_person_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_deposit_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_debitcard_1.parquet\", 1),\n    ],\n    \"depth_2\": [\n        read_file(TRAIN_DIR / \"train_credit_bureau_b_2.parquet\", 2),\n    ]\n}","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:17:19.328228Z","iopub.execute_input":"2024-05-31T06:17:19.328717Z","iopub.status.idle":"2024-05-31T06:17:51.782024Z","shell.execute_reply.started":"2024-05-31T06:17:19.328688Z","shell.execute_reply":"2024-05-31T06:17:51.781194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Merge the training dataset file\ndf_train = feature_eng(**data_store)\nprint(\"train data shape:\\t\", df_train.shape)","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:17:51.783125Z","iopub.execute_input":"2024-05-31T06:17:51.783415Z","iopub.status.idle":"2024-05-31T06:17:57.140681Z","shell.execute_reply.started":"2024-05-31T06:17:51.783392Z","shell.execute_reply":"2024-05-31T06:17:57.139868Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Read the testing dataset file\ndata_store = {\n    \"df_base\": read_file(TEST_DIR / \"test_base.parquet\"),\n    \"depth_0\": [\n        read_file(TEST_DIR / \"test_static_cb_0.parquet\"),\n        read_files(TEST_DIR / \"test_static_0_*.parquet\"),\n    ],\n    \"depth_1\": [\n        read_files(TEST_DIR / \"test_applprev_1_*.parquet\", 1),\n        read_file(TEST_DIR / \"test_tax_registry_a_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_tax_registry_b_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_tax_registry_c_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_credit_bureau_b_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_other_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_person_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_deposit_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_debitcard_1.parquet\", 1),\n    ],\n    \"depth_2\": [\n        read_file(TEST_DIR / \"test_credit_bureau_b_2.parquet\", 2),\n    ]\n}","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:17:57.141774Z","iopub.execute_input":"2024-05-31T06:17:57.142064Z","iopub.status.idle":"2024-05-31T06:17:57.539949Z","shell.execute_reply.started":"2024-05-31T06:17:57.142039Z","shell.execute_reply":"2024-05-31T06:17:57.539148Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Merge the testing dataset file\ndf_test = feature_eng(**data_store)\nprint(\"test data shape:\\t\", df_test.shape)","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:17:57.541139Z","iopub.execute_input":"2024-05-31T06:17:57.541487Z","iopub.status.idle":"2024-05-31T06:17:57.569620Z","shell.execute_reply.started":"2024-05-31T06:17:57.541458Z","shell.execute_reply":"2024-05-31T06:17:57.568840Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Filter\ndf_train = df_train.pipe(Pipeline.filter_cols)\ndf_test = df_test.select([col for col in df_train.columns if col != \"target\"])\n\nprint(\"train data shape:\\t\", df_train.shape)\nprint(\"test data shape:\\t\", df_test.shape)","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:17:57.570540Z","iopub.execute_input":"2024-05-31T06:17:57.570798Z","iopub.status.idle":"2024-05-31T06:17:59.466146Z","shell.execute_reply.started":"2024-05-31T06:17:57.570777Z","shell.execute_reply":"2024-05-31T06:17:59.465156Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Convert from polar to pandas\ndf_train, cat_cols = to_pandas(df_train)\ndf_test, cat_cols = to_pandas(df_test, cat_cols)","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:17:59.467548Z","iopub.execute_input":"2024-05-31T06:17:59.468112Z","iopub.status.idle":"2024-05-31T06:18:14.503070Z","shell.execute_reply.started":"2024-05-31T06:17:59.468079Z","shell.execute_reply":"2024-05-31T06:18:14.502225Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.head()","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:18:14.508267Z","iopub.execute_input":"2024-05-31T06:18:14.508844Z","iopub.status.idle":"2024-05-31T06:18:14.536543Z","shell.execute_reply.started":"2024-05-31T06:18:14.508818Z","shell.execute_reply":"2024-05-31T06:18:14.535658Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# To check \"null\" values\ndf_train.isnull().any()","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:18:14.537911Z","iopub.execute_input":"2024-05-31T06:18:14.538147Z","iopub.status.idle":"2024-05-31T06:18:14.868537Z","shell.execute_reply.started":"2024-05-31T06:18:14.538127Z","shell.execute_reply":"2024-05-31T06:18:14.867527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.describe()","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:18:14.869478Z","iopub.execute_input":"2024-05-31T06:18:14.870205Z","iopub.status.idle":"2024-05-31T06:18:28.095284Z","shell.execute_reply.started":"2024-05-31T06:18:14.870172Z","shell.execute_reply":"2024-05-31T06:18:28.094359Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.info()","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:18:28.096431Z","iopub.execute_input":"2024-05-31T06:18:28.096954Z","iopub.status.idle":"2024-05-31T06:18:28.134897Z","shell.execute_reply.started":"2024-05-31T06:18:28.096926Z","shell.execute_reply":"2024-05-31T06:18:28.133901Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Find out how many null values in the dataset\ndf_train.isnull().sum().sum()","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:18:28.136079Z","iopub.execute_input":"2024-05-31T06:18:28.136786Z","iopub.status.idle":"2024-05-31T06:18:28.665934Z","shell.execute_reply.started":"2024-05-31T06:18:28.136761Z","shell.execute_reply":"2024-05-31T06:18:28.665030Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\n\ndef clear_columns_with_nan(df):\n    \"\"\"\n    Remove columns from the DataFrame that contain at least one NaN value.\n\n    Parameters:\n    df (pd.DataFrame): The input DataFrame.\n\n    Returns:\n    pd.DataFrame: The DataFrame with columns containing NaN values removed.\n    \"\"\"\n    # Check for total missing values in the dataset\n    total_missing = df.isnull().sum().sum()\n    print(\"\\nTotal Missing values in dataset: \" + str(total_missing))\n    \n    # If there are missing values, drop columns with any NaN values\n    if total_missing > 0:\n        df = df.dropna(axis=1, how='any')\n    \n    return df\n\n# Clear columns with NaN values if any missing values are detected\ndf_cleaned = clear_columns_with_nan(df_train)\n\n# Print cleaned DataFrame\nprint(\"\\nDataFrame after clearing columns with NaN values:\")\nprint(df_cleaned)","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2024-05-31T06:18:28.667401Z","iopub.execute_input":"2024-05-31T06:18:28.667785Z","iopub.status.idle":"2024-05-31T06:18:30.010319Z","shell.execute_reply.started":"2024-05-31T06:18:28.667754Z","shell.execute_reply":"2024-05-31T06:18:30.009372Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_cleaned.isnull().any()\n# There is no null values after the df_train has been cleaned","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:18:30.011490Z","iopub.execute_input":"2024-05-31T06:18:30.011809Z","iopub.status.idle":"2024-05-31T06:18:30.104970Z","shell.execute_reply.started":"2024-05-31T06:18:30.011783Z","shell.execute_reply":"2024-05-31T06:18:30.103876Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_cleaned.info()","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:18:30.106305Z","iopub.execute_input":"2024-05-31T06:18:30.107215Z","iopub.status.idle":"2024-05-31T06:18:30.286001Z","shell.execute_reply.started":"2024-05-31T06:18:30.107182Z","shell.execute_reply":"2024-05-31T06:18:30.285233Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <p style =\"color: blue;\">3. Exploratory Data Analysis</p>","metadata":{}},{"cell_type":"markdown","source":"Exploratory data analysis is an approach of analyzing data sets to summarize their main characteristics, often using statistical graphics and other data visualization methods. \n### Description of dataframe\nIn this section, there are 8 features that selected from the dataset that has been cleaned to use. \n- case_id\n- WEEK_NUM\n- target\n- month_decision\n- weekday_decision\n- annuity_780A\n- max_mainoccupationinc_384A\n- max_incometype_1044T \n\n### Graphical plots of data\nSkewness Graph: To determine the symmetry of distribution (monthly_annuity & max_income is positively skewed)\n\nBoxplot Graph: To determine the outliers (monthly_annuity & max_income do have outliers)\n\n### Descriptive statistics of data\nAfter reading all the csv files, the total number of rows and columns of all the data in each dataframe is shown by using the df.shape function. It can be seen that df_eda has 1526659 rows and 8 columns, df_test has 10 rows and 266 columns and df_train has 1526659 rows and 267 columns. The total number of missing data in df_train is shown by the df_train.isnull().sum() function .","metadata":{}},{"cell_type":"code","source":"df_eda = pd.DataFrame()\n\ndf_eda[\"case_id\"] = df_cleaned[\"case_id\"]\ndf_eda['week_num'] = df_cleaned['WEEK_NUM']\ndf_eda['target'] = np.where(df_cleaned['target']< 1, \"no\", \"yes\")\ndf_eda['month_decision'] = df_cleaned['month_decision']\ndf_eda['weekday_decision'] = df_cleaned['weekday_decision']\ndf_eda['monthly_annuity'] = df_cleaned['annuity_780A']\ndf_eda['max_income'] = df_cleaned['max_mainoccupationinc_384A']\ndf_eda['income_type'] = df_cleaned['max_incometype_1044T']\n\n# df_eda.head().style.background_gradient(cmap = \"Reds\").set_properties(**{\"font-family\" : \"Segoe UI\"}).hide_index()","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:18:30.286842Z","iopub.execute_input":"2024-05-31T06:18:30.287234Z","iopub.status.idle":"2024-05-31T06:18:30.540410Z","shell.execute_reply.started":"2024-05-31T06:18:30.287210Z","shell.execute_reply":"2024-05-31T06:18:30.539671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check dimensions of the array\nprint(df_train.shape)\nprint(df_test.shape)\nprint(df_eda.shape)","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:18:30.541330Z","iopub.execute_input":"2024-05-31T06:18:30.541819Z","iopub.status.idle":"2024-05-31T06:18:30.547316Z","shell.execute_reply.started":"2024-05-31T06:18:30.541784Z","shell.execute_reply":"2024-05-31T06:18:30.546321Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# This is to list the numerical and categorical features in the dataset\ncol = list(df_eda.columns)\ncategorical_features = []\nnumerical_features = []\nfor i in col:\n    if len(df_eda[i].unique()) > 6:\n        numerical_features.append(i)\n    else:\n        categorical_features.append(i)\n\nprint('Categorical Features :',*categorical_features)\nprint('Numerical Features :',*numerical_features)","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:18:30.548713Z","iopub.execute_input":"2024-05-31T06:18:30.549240Z","iopub.status.idle":"2024-05-31T06:18:30.816334Z","shell.execute_reply.started":"2024-05-31T06:18:30.549214Z","shell.execute_reply":"2024-05-31T06:18:30.815238Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate skewness for numerical columns\nskewness = df_eda.select_dtypes(include=['int64', 'float64']).skew()\n\n# Count the number of numerical columns\nnum_cols_count = len(df_eda.select_dtypes(include=['int64', 'float64']).columns)\n\n# Determine the layout for subplots\nnum_rows = (num_cols_count + 3) // 4  # Adjust the number of columns in each row\nnum_cols = min(4, num_cols_count)  # Maximum of 4 columns in each row\n\n# Plot histograms for numerical columns to visualize distributions and identify anomalies\nfig, axes = plt.subplots(num_rows, num_cols, figsize=(20, num_rows * 5))  # Increased figure size\n\n# Flatten axes array for easy iteration\naxes = axes.flatten()\n\nfor col_idx in range(num_cols_count):\n    col = df_eda.select_dtypes(include=['int64', 'float64']).columns[col_idx]\n    axes[col_idx].hist(df_eda[col], bins=15, color='green', alpha=0.7, edgecolor='black')  # Added edgecolor\n    axes[col_idx].set_title(f'{col}')\n    axes[col_idx].set_xlabel(col)\n    axes[col_idx].set_ylabel('Frequency')\n    skew_val = skewness[col]\n    axes[col_idx].text(0.95, 0.95, f'Skewness: {skew_val:.2f}', horizontalalignment='right',\n                       verticalalignment='top', transform=axes[col_idx].transAxes, fontsize=10, color='red')\n\n# Hide any unused subplots\nfor i in range(num_cols_count, len(axes)):\n    fig.delaxes(axes[i])\n\nplt.tight_layout()\nplt.show()\n\n# Print skewness values\nprint(\"Skewness:\")\nprint(skewness)\n","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:18:30.817583Z","iopub.execute_input":"2024-05-31T06:18:30.817943Z","iopub.status.idle":"2024-05-31T06:18:32.417344Z","shell.execute_reply.started":"2024-05-31T06:18:30.817916Z","shell.execute_reply":"2024-05-31T06:18:32.416369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Filter numeric columns\nnumeric_cols = df_eda.select_dtypes(include=['int64', 'float64']).columns\n\n# Plotting boxplots for each numerical feature to identify outliers\nfor column in numeric_cols:\n    plt.figure(figsize=(10, 6))\n    sns.boxplot(x=df_eda[column],palette='rainbow')\n    plt.title(f'Boxplot of {column}')\n    plt.show()","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2024-05-31T06:18:32.418598Z","iopub.execute_input":"2024-05-31T06:18:32.419367Z","iopub.status.idle":"2024-05-31T06:18:33.652271Z","shell.execute_reply.started":"2024-05-31T06:18:32.419335Z","shell.execute_reply":"2024-05-31T06:18:33.651336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <p style =\"color: blue;\">4. Model Building</p>","metadata":{}},{"cell_type":"markdown","source":"#### (1) How the data is partitioned\nTo begin the model building process, we split the df_eda into a training and testing dataset (80% Training, 20% Test). And this is easily done using the sci-kit learn module. The features encoded using one-hot encoding to handle categorical variables. \n\n#### (2) How the model is chosen\nThe performance of models using metrics such as Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), and R-squared (R2). The model with the lowest test RMSE is selected as the best model. \n\n#### (3) How the model is trained\nBefore fitting the data into the model, all the current data has to be in numerical form. After changing the data, we could then proceed with fitting the model to our dataset.\n\n#### (4) The attributes that have the greatest effect on the results","metadata":{}},{"cell_type":"code","source":"x = df_eda.drop('target',axis=1)\ny = df_eda['target']\n\nX_encoded = pd.get_dummies(x, drop_first=True)\n\nx_train,x_test,y_train,y_test = train_test_split(X_encoded,y,test_size = 0.15,shuffle = True,random_state=42)","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:18:33.653501Z","iopub.execute_input":"2024-05-31T06:18:33.653861Z","iopub.status.idle":"2024-05-31T06:18:34.033323Z","shell.execute_reply.started":"2024-05-31T06:18:33.653826Z","shell.execute_reply":"2024-05-31T06:18:34.032541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define a function to evaluate the models using metrics like Mean Absolute Error (MAE), \n# Root Mean Squared Error (RMSE), and R-squared (R2).\ndef evaluate_model(y_true, y_pred):\n    from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score\n    mae = mean_absolute_error(y_true, y_pred)\n    rmse = mean_squared_error(y_true, y_pred, squared=False)\n    r2 = r2_score(y_true, y_pred)\n    return mae, rmse, r2","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:18:34.034451Z","iopub.execute_input":"2024-05-31T06:18:34.034867Z","iopub.status.idle":"2024-05-31T06:18:34.039967Z","shell.execute_reply.started":"2024-05-31T06:18:34.034842Z","shell.execute_reply":"2024-05-31T06:18:34.039174Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.preprocessing import LabelEncoder\n\n# build a model\nmodels = {\n    \"Logistic Regression\": LogisticRegression(),\n    \"Decision Tree\": DecisionTreeRegressor(),\n}\n\n# Use LabelEncoder to convert categorical target variables (if any) to numerical values.\nlabel_encoder = LabelEncoder()\ny_train_encoded = label_encoder.fit_transform(y_train)\ny_test_encoded = label_encoder.transform(y_test)","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:18:34.041300Z","iopub.execute_input":"2024-05-31T06:18:34.041576Z","iopub.status.idle":"2024-05-31T06:18:34.824343Z","shell.execute_reply.started":"2024-05-31T06:18:34.041552Z","shell.execute_reply":"2024-05-31T06:18:34.823350Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <p style =\"color: blue;\">5. Model Evaluation</p>","metadata":{}},{"cell_type":"markdown","source":"### Describe the process of model evaluation\n#### The performance of the model created\nModel 1 | Logistic Regression\n\nThis model using logistic regression scored an even higher accuracy of 96%. Both training and testing performances are quite similar, indicating the model is not overfitting. However, the R2 score is negative, suggesting that the model performs worse than a horizontal line. This might indicate that the model is not capturing the variance in the data well.\n\nModel 2 | Deicision Tree\n\nThe perfect R2 score on the training set and a significantly lower R2 score on the testing set suggest overfitting. The model performs poorly on the testing set, indicating that it fails to generalize well to unseen data.\n\n#### How the model can be used to produce stable score\nLogistic Regression seems to be a better choice compared to Decision Tree based on these results.\nFurther analysis and feature engineering might be needed to improve the performance of the Logistic Regression model. Consider exploring feature importance, tuning hyperparameters, or trying different algorithms.\nSince the R2 scores for both models are negative or extremely low, it's important to reassess the features and model selection process to improve model performance.","metadata":{"vscode":{"languageId":"raw"}}},{"cell_type":"code","source":"# Fit models and make predictions\nfor name, model in models.items():\n    model.fit(x_train, y_train_encoded)\n    \n    # Make prediction:\n    y_train_pred = model.predict(x_train)\n    y_test_pred = model.predict(x_test)\n    \n    # Evaluate Train and Test dataset :\n    model_train_mae, model_train_rmse, model_train_r2 = evaluate_model(y_train_encoded, y_train_pred)\n    model_test_mae, model_test_rmse, model_test_r2 = evaluate_model(y_test_encoded, y_test_pred)\n    \n    print(name)\n    print(\"Model Performance for Training set:\")\n    print('Root Mean Squared Error:', model_train_rmse)\n    print(\"Mean Absolute Error:\", model_train_mae)\n    print(\"R2 Score:\", model_train_r2)\n    print(\"----------------------------------------------------\")\n    print(\"Model Performance for Testing set:\")\n    print('Root Mean Squared Error:', model_test_rmse)\n    print('Mean Absolute Error:', model_test_mae)\n    print('R2 Score:', model_test_r2)\n    print()","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2024-05-31T06:18:34.825749Z","iopub.execute_input":"2024-05-31T06:18:34.826546Z","iopub.status.idle":"2024-05-31T06:18:53.636762Z","shell.execute_reply.started":"2024-05-31T06:18:34.826518Z","shell.execute_reply":"2024-05-31T06:18:53.635763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Initialize a list to store model performance metrics\nmodel_performance = []\n\n# Fit models and make predictions\nfor name, model in models.items():\n    model.fit(x_train, y_train_encoded)\n    \n    # Make predictions\n    y_train_pred = model.predict(x_train)\n    y_test_pred = model.predict(x_test)\n    \n    # Evaluate Train and Test datasets\n    model_train_mae, model_train_rmse, model_train_r2 = evaluate_model(y_train_encoded, y_train_pred)\n    model_test_mae, model_test_rmse, model_test_r2 = evaluate_model(y_test_encoded, y_test_pred)\n    \n    # Store the metrics in the list\n    model_performance.append({\n        'Model': name,\n        'Dataset': 'Train',\n        'RMSE': model_train_rmse,\n        'MAE': model_train_mae,\n        'R2': model_train_r2\n    })\n    model_performance.append({\n        'Model': name,\n        'Dataset': 'Test',\n        'RMSE': model_test_rmse,\n        'MAE': model_test_mae,\n        'R2': model_test_r2\n    })\n\n# Convert the list of dictionaries to a DataFrame\nperformance_df = pd.DataFrame(model_performance)\n\n# Plotting the performance metrics\nfig, axes = plt.subplots(1, 3, figsize=(18, 6))\n\n# RMSE Plot\nsns.barplot(data=performance_df, x='Model', y='RMSE', hue='Dataset', ax=axes[0])\naxes[0].set_title('RMSE by Model')\naxes[0].set_xticklabels(axes[0].get_xticklabels(), rotation=45)\n\n# MAE Plot\nsns.barplot(data=performance_df, x='Model', y='MAE', hue='Dataset', ax=axes[1])\naxes[1].set_title('MAE by Model')\naxes[1].set_xticklabels(axes[1].get_xticklabels(), rotation=45)\n\n# R2 Plot\nsns.barplot(data=performance_df, x='Model', y='R2', hue='Dataset', ax=axes[2])\naxes[2].set_title('R2 Score by Model')\naxes[2].set_xticklabels(axes[2].get_xticklabels(), rotation=45)\n\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:18:53.638003Z","iopub.execute_input":"2024-05-31T06:18:53.638361Z","iopub.status.idle":"2024-05-31T06:19:14.268361Z","shell.execute_reply.started":"2024-05-31T06:18:53.638311Z","shell.execute_reply":"2024-05-31T06:19:14.267415Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <p style =\"color: blue;\">6. Model Validation</p>","metadata":{}},{"cell_type":"markdown","source":"While model validation and model evaluation may seem like very similar things, model validation is the act of ensuring models are performing as they should, while model evaluation is an assessment. In other words, we validate our models to test their accuracies on the data we already have. The most common model validation we have can be seen in our evaluation above - the Train/test split method. \n\nIn this section, we will use another cross validation method called K-fold cross validation.","metadata":{"vscode":{"languageId":"raw"}}},{"cell_type":"code","source":"pairs_selectedfeatures = ['case_id', 'WEEK_NUM', 'target', 'month_decision', 'weekday_decision', 'annuity_780A', 'max_mainoccupationinc_384A', 'max_incometype_1044T']\n\n#X = df_test[pairs_selectedfeatures]\n#y = df_test['target']\n\nX_test = df_test.drop(columns=[\"WEEK_NUM\"])\nX_test = X_test.set_index(\"case_id\")","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:19:14.269805Z","iopub.execute_input":"2024-05-31T06:19:14.270159Z","iopub.status.idle":"2024-05-31T06:19:14.280111Z","shell.execute_reply.started":"2024-05-31T06:19:14.270132Z","shell.execute_reply":"2024-05-31T06:19:14.279108Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#loading the cross validation | K-fold method\n#splits into 5 sets of data\nfrom sklearn.datasets import load_iris\nfrom sklearn.model_selection import cross_val_score, KFold\nfrom sklearn.linear_model import LogisticRegression\n\n# Load the Iris dataset\niris = load_iris()\niris_X = iris.data\niris_y = iris.target\n\n# Initialize the logistic regression model\nlogreg = LogisticRegression()\n\n# Initialize KFold cross-validation with 5 splits\nkf = KFold(n_splits=5)\n\n# Perform cross-validation and compute scores\nscores = cross_val_score(logreg, iris_X, iris_y, cv=kf)\n\n# Print the cross-validation scores\nprint(\"Cross Validation Scores:\", scores)\nprint(\"Average Cross Validation Score:\", scores.mean())\nprint(\"Standard Deviation of Cross Validation Scores:\", scores.std())","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:19:14.281690Z","iopub.execute_input":"2024-05-31T06:19:14.282059Z","iopub.status.idle":"2024-05-31T06:19:14.465594Z","shell.execute_reply.started":"2024-05-31T06:19:14.281968Z","shell.execute_reply":"2024-05-31T06:19:14.464646Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Ensure df_cleaned and pairs_selectedfeatures are defined\n# One-hot encode the 'income_type' column\n#X_encoded = pd.get_dummies(df_cleaned[pairs_selectedfeatures], drop_first=True)\n\n# Initialize and fit the Logistic Regression model with training data\n#logreg = LogisticRegression()\n#logreg.fit(X_encoded, y)\n\n# Make predictions on df_cleaned and extract probabilities of the positive class\n#y_pred = pd.Series(logreg.predict_proba(X_test)[:, 1], index=df_cleaned.index)\n\n# Load the sample submission file\n#df_subm = pd.read_csv(ROOT / \"sample_submission.csv\")\n\n# Ensure the case_id in the sample submission matches the index of predictions\n#if not df_subm['case_id'].isin(y_pred.index).all():\n#    raise ValueError(\"Mismatch between 'case_id' in submission file and index of predictions\")\n\n# Set 'case_id' as the index of the submission DataFrame\n#df_subm = df_subm.set_index(\"case_id\")\n\n# Check for missing values and fill if necessary\n#df_subm['score'] = df_subm['score'].fillna(0)\n\n# Assign predicted probabilities to the 'score' column in the submission DataFrame\n#df_subm[\"score\"] = df_subm.index.to_series().map(y_pred)\n\n# Check for any missing values in the 'score' column\n#if df_subm['score'].isnull().any():\n#    raise ValueError(\"There are missing values in the 'score' column after assignment\")\n\n# Save the submission DataFrame to a CSV file\n#df_subm.to_csv(\"submission.csv\")","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2024-05-31T06:19:14.466604Z","iopub.execute_input":"2024-05-31T06:19:14.466883Z","iopub.status.idle":"2024-05-31T06:19:14.473504Z","shell.execute_reply.started":"2024-05-31T06:19:14.466858Z","shell.execute_reply":"2024-05-31T06:19:14.472629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.impute import SimpleImputer\nfrom sklearn.preprocessing import OneHotEncoder\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.pipeline import Pipeline\n\n# List of selected attributes\nselected_attributes = ['case_id', 'WEEK_NUM', 'month_decision', 'weekday_decision', 'annuity_780A', \n                       'max_mainoccupationinc_384A', 'max_incometype_1044T']\n\n# Extract numerical and categorical attributes\nnumerical_attributes = ['WEEK_NUM', 'month_decision', 'weekday_decision', 'annuity_780A', \n                        'max_mainoccupationinc_384A']\ncategorical_attributes = ['max_incometype_1044T']\n\n# Define preprocessing steps for numerical and categorical data\nnumerical_transformer = SimpleImputer(strategy='mean')\ncategorical_transformer = Pipeline(steps=[\n    ('imputer', SimpleImputer(strategy='most_frequent')),\n    ('onehot', OneHotEncoder(handle_unknown='ignore'))\n])\n\n# Combine preprocessing steps\npreprocessor = ColumnTransformer(\n    transformers=[\n        ('num', numerical_transformer, numerical_attributes),\n        ('cat', categorical_transformer, categorical_attributes)\n    ])\n\n# Define the model\nmodel = Pipeline(steps=[('preprocessor', preprocessor),\n                        ('classifier', LogisticRegression())])\n\n# Train the model\nX_train = df_train[selected_attributes]\ny_train = df_train['target']\nmodel.fit(X_train, y_train)\n\n# Make predictions on the testing data\nX_test = df_test[selected_attributes]\ntest_predictions = model.predict_proba(X_test)[:, 1]\n\n# Create a DataFrame for submission\nsubmission_df = pd.DataFrame({'case_id': df_test['case_id'], 'score': test_predictions})\n\n# Save the submission DataFrame to a CSV file\nsubmission_df.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:19:14.474583Z","iopub.execute_input":"2024-05-31T06:19:14.474842Z","iopub.status.idle":"2024-05-31T06:19:26.096794Z","shell.execute_reply.started":"2024-05-31T06:19:14.474819Z","shell.execute_reply":"2024-05-31T06:19:26.095465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#submission = pd.DataFrame({\n#    \"case_id\": df_cleaned[\"case_id\"].to_numpy(),\n#    \"score\":y_pred\n#}).set_index('case_id')\n\n#submission.to_csv(\"./submission.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:19:26.098680Z","iopub.execute_input":"2024-05-31T06:19:26.099328Z","iopub.status.idle":"2024-05-31T06:19:26.103713Z","shell.execute_reply.started":"2024-05-31T06:19:26.099296Z","shell.execute_reply":"2024-05-31T06:19:26.102842Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_subm.head()","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:19:26.105122Z","iopub.execute_input":"2024-05-31T06:19:26.105747Z","iopub.status.idle":"2024-05-31T06:19:26.113417Z","shell.execute_reply.started":"2024-05-31T06:19:26.105717Z","shell.execute_reply":"2024-05-31T06:19:26.112534Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_subm.dtypes","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:19:26.114821Z","iopub.execute_input":"2024-05-31T06:19:26.115292Z","iopub.status.idle":"2024-05-31T06:19:26.122450Z","shell.execute_reply.started":"2024-05-31T06:19:26.115261Z","shell.execute_reply":"2024-05-31T06:19:26.121576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#print(\"Check null: \", df_subm[\"score\"].isnull().any())","metadata":{"execution":{"iopub.status.busy":"2024-05-31T06:19:26.123965Z","iopub.execute_input":"2024-05-31T06:19:26.125147Z","iopub.status.idle":"2024-05-31T06:19:26.130947Z","shell.execute_reply.started":"2024-05-31T06:19:26.125115Z","shell.execute_reply":"2024-05-31T06:19:26.130012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Explanation\nThe chosen validation approach is k-Fold Cross Validation, specifically a 5-fold approach where the dataset is randomly shuffled and divided into 5 groups. Each group serves as a test set once, while the remaining 4 act as training sets. After each iteration, the average score across all folds is computed.\n\nIn terms of model performance, the logistic regression model achieved an accuracy of 0.93, with a standard deviation of 0.068. This is slightly lower compared to the accuracy of 0.96 obtained in a previous analysis. \nThe cross-validation scores range from 0.8333 to 1.0, indicating that the model performs well across different folds.\nThe average cross-validation score is approximately 0.9267, suggesting good overall performance.\nThe standard deviation of the scores is relatively small (approximately 0.0680), indicating consistency in model performance across different folds.","metadata":{"vscode":{"languageId":"raw"}}},{"cell_type":"markdown","source":"## <p style =\"color: blue;\">7. Discussion</p>","metadata":{}},{"cell_type":"markdown","source":"### Identify\n#### (1) The factors that have significant influences on the credit risk and its stability\nThe selected variables for exploratory data analysis (EDA) in this dataset appear suitable for investigating credit risk and its stability. Here's an overview:\n\n- case_id: Unique identifier for each loan application, essential for tracking.\n- WEEK_NUM: Week of the year of the application, useful for identifying weekly patterns.\n- target: Indicates loan repayment status, crucial for model training.\n- month_decision: Month of loan decision, allows for seasonal trend analysis.\n- weekday_decision: Day of the week of the decision, useful for weekly pattern analysis.\n- annuity_780A: Loan annuity amount, indicates repayment capacity.\n- max_mainoccupationinc_384A: Maximum income from main occupation, reflects financial stability.\n- max_incometype_1044T: Maximum income from all sources, provides a comprehensive income view.\n\nThese variables offer a mix of temporal, financial, and categorical data critical for understanding and predicting credit risk.\n\n#### (2) Any interesting observation from this challenge\nFrom this challenge, we know that the score is a immportant information to determine which clients are more likely to default on their loans. By knowing score, we know that which client have low or high loan risk. We use income type and income amount to calculate the score in order to predict the loan risk.","metadata":{"tags":[],"vscode":{"languageId":"raw"}}},{"cell_type":"markdown","source":"### Explain\n#### (1) The limitations and weaknesses of the modelling approach\n- Logistic Regression\n\nThe negative R2 score indicates that the model is performing worse than a horizontal line (a baseline model that simply predicts the mean of the target variable). This suggests that Logistic Regression is not capturing the underlying patterns in the data effectively.\n\nMean Absolute Error (MAE) and Root Mean Squared Error (RMSE) are regression metrics, which are not suitable for evaluating classification models like Logistic Regression. Classification models should be evaluated using metrics such as accuracy, precision, recall, F1 score, ROC-AUC, etc.\n\nIf the target variable is imbalanced (many more \"no\" than \"yes\" cases), Logistic Regression might not perform well without techniques to handle imbalance, such as oversampling the minority class, undersampling the majority class, or using class weights.\n\n- Decision Tree\n\nThe Decision Tree model shows perfect performance on the training set (RMSE of 0.0 and R2 of 1.0) but poor performance on the testing set (higher RMSE and negative R2 score). This is a clear sign of overfitting, where the model has learned the training data too well, including noise and outliers, and fails to generalize to new, unseen data.\n\n\n#### (2) The steps taken to improve the matching accuracy in your modelling approach\nFrom the merged and cleaned dataset, we have chosen only a few attributes to be used and dropped the attributes that are of no use. This will increase the accuracy as there are less missing values.","metadata":{}},{"cell_type":"markdown","source":"### Elaborate\n#### (1) The experience in participating in a Kaggle challenge\nWe can gain more knowledge about the practical applications of data science by taking part in this Kaggle competition. To help create a stronger model, the team put a lot of effort into learning the Kaggle datasets and researching related topics online. It was a great task to have, and we were able to get to know each other better because of our frequent communication.\n#### (2) The discussion and submission score on Kaggle (include the screenshot or link to your submission here)\nProof of discussion/participation : [Link]()\n#### (3) The improvements that need to be done in order to win the challenge\nThe improvements that could be done would be to build a better model with higher accuracy, to acheive this, the team would have to utlilize more of the data from the original dataset, followed by intensive data pre-processing (such as fixing all the categories into a proper format, translating any non-english texts, and so on). Doing this would have led to a more accurate & precise model.","metadata":{}},{"cell_type":"markdown","source":"# Team contribution","metadata":{}},{"cell_type":"markdown","source":"| Name | Tasks Contributed | Contribution Percentage Overall Submission |\n| --- | --- | --- |\n| Nicholas Ting Tiew Leong (102767857) | Documentation, EDA, Model Building, Model Evaluation, Cross Validation | 50 |\n| Ng Ye Ding (101234017) | Documentation, EDA, Model Building, Model Evaluation, Cross Validation | 50 |","metadata":{"vscode":{"languageId":"raw"}}}]}