{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":50160,"databundleVersionId":7921029,"sourceType":"competition"}],"dockerImageVersionId":30664,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Created by yunsuxiaozi 2024/3/12\n\n#### It has to be said that there are many documents for this competition.\n\n#### In the days leading up to the suspension of the competition, I was still trying to integrate more files to construct more features to improve my score. Currently, I have not directly copied any open-source high scoring code. This is partly due to the existence of \"shakes\" in data mining competitions, and I also want to score on my own. The problem with this is that I spent a lot of time \"deal with threw exceptions\" in this competition.\n\n#### So far, I haven't fully integrated all the files, and haven't even looked carefully at what each file specifically contains. Therefore, today I plan to slow down, carefully study the content of the competition files, and provide some unique insights of my own.\n\n#### The current version may not be able to view all file contents, and updates will be completed in the future.","metadata":{}},{"cell_type":"code","source":"import pandas as pd#导入csv文件的库\nimport time#标准库的时间模块\n#为了方便后期调用训练的模型时不会调用错版本,提供模型训练的时间\n#time.strftime()函数用于将时间对象格式化为字符串，time.localtime()函数返回表示当前本地时间的time.struct_time对象\ncurrent_time = time.strftime(\"%Y-%m-%d %H:%M:%S\", time.localtime())\nprint(\"this notebook training time is \", current_time)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## base.csv\n\n- 'case_id':I am currently just treating it as a regular id.\n- 'date_decision','MONTH','WEEK_NUM':I am currently not using them as input variables for the model.\n- 'target': This is the label for the binary classification task.","metadata":{}},{"cell_type":"code","source":"#case_id是目前只当它是普通id,date_decision,month,week_num这些时间特征不考虑使用.target有点用.\nbase=pd.read_csv(\"/kaggle/input/home-credit-credit-risk-model-stability/csv_files/train/train_base.csv\")\nbase","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## feature_definitions.csv \n\nWe will provide official definitions for the column names of each file next.","metadata":{}},{"cell_type":"code","source":"feature=pd.read_csv(\"/kaggle/input/home-credit-credit-risk-model-stability/feature_definitions.csv\")\nprint(f\"len(feature):{len(feature)}\")\nfeature","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## applprev_1.csv (1_0.csv,1_1.csv)\n\n### Due to the identical column names of the two lists, only one list is opened here.","metadata":{}},{"cell_type":"code","source":"applprev1=pd.read_csv(\"/kaggle/input/home-credit-credit-risk-model-stability/csv_files/train/train_applprev_1_0.csv\")\napplprev1","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tmp=feature[feature['Variable'].isin(applprev1.columns)]\nfor idx in range(len(tmp)):\n    print(tmp['Variable'].values[idx],tmp['Description'].values[idx])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Variable interpretation and suggestion:\n\n- 'actualdpd_943P':This is the number of days overdue from the previous contract, the more days there are, the more severe it becomes.If the overdue days are 0, they may need to be separately classified as Class 1, or the days can be classified into different levels. For filling in missing values, \"ffill\" can be considered.\n\n- 'annuity_853A':Monthly annuity for previous applications.This variable can be considered using conventional statistical modeling methods.\n\n- 'approvaldate_319D':Approval Date of Previous Application.Missing values may not have been applied for before, and the date column can be divided into two categories.\n\n- 'byoccupationinc_3656910L':Applicant's income from previous applications.The income level is related to whether there is a default, and filling in missing values here can be done using the median.\n\n- 'cancelreason_3545846M':Application cancellation reason.There are many categories here, except for 'a55475b1', which are all strings with the shape of 'P94_109_143'. Using one-hot encoding here will consume a lot of memory, so consider using the first letter as the one-hot encoding.\n\n- 'childnum_21L':How many children were there in the last application.Considering that first-time users of credit cards may be young or have no money, fill in the missing value as 0.\n\n- 'creationdate_885D':Date when previous application was created.This can be compared to the approval date of the previous application. If the approval time is too long, may he have a breach of contract record?\n\n- 'credacc_actualbalance_314A':Actual balance on credit account.This balance needs to be updated.\n\n- 'credacc_credlmt_575A':Credit card credit limit provided for previous applications.\n\n- 'credacc_maxhisbal_375A':Maximal historical balance of previous credit account.The more balance there is, the less likely it is to default?\n\n- 'credacc_minhisbal_90A':Minimum historical balance of previous credit accounts.If the balance is negative, there is a possibility of default.\n\n- 'credacc_status_367L':Account status of previous credit applications.Here,we can consider one-hot encoding.\n\n- 'credacc_transactions_402L':Number of transactions made with the previous credit account of the applicant.We need everyone's newest data here.\n\n- 'credamount_590A':Loan amount or card limit of previous applications.\n\n- 'credtype_587L':Credit type of previous application.Here we use regular one hot encoding.\n\n- 'currdebt_94A':Previous application's current debt.People without debt are even less likely to default, and we also need to obtain the latest records here.\n\n- 'dateactivated_425D':Contract activation date of the applicant's previous application.This may not be of much use?\n\n- 'district_544M':District of the address used in the previous loan application.May we consider the wealth gap in each region?\n\n- 'downpmt_134A': Previous application downpayment amount.numerical variable.\n\n- 'dtlastpmt_581D':Date of last payment made by the applicant.This may not be of much use?\n\n- 'dtlastpmtallstes_3545839D':Date of the applicant's last payment.Is there any difference between this and the previous feature?If these two columns are not missing values and there are still some different data, why is this?\n\n- 'education_1138M':Applicant's education level from their previous application.However,this information seems to have been encrypted, so it can only be encoded using one-hot-encoding.\n\n\n- 'employedfrom_700D':Employment start date from the previous application.Here, we can consider the number of people in each state or perform one hot encoding.\n\n- 'familystate_726L':Family State in previous application of applicant.We can consider living alone or using simple one hot encoding here.\n\n- 'firstnonzeroinstldate_307D':Date of first instalment in the previous application.\n\n- 'inittransactioncode_279L':Type of the initial transaction made in the previous application of the client.Here we can use one-hot-encoding.\n\n- 'isbidproduct_390L':Flag for determining if the product is a cross-sell in previous applications.False accounts for the majority, and missing values can be filled with False.\n\n- 'isdebitcard_527L':Previous application flag indicating if product being applied for is a debit card.False accounts for the majority, and missing values can be filled with False.\n\n- 'mainoccupationinc_437A':Client's main income amount in their previous application.\n\n- 'maxdpdtolerance_577P':Maximum DPD with tolerance (on previous application/s).\n\n- 'outstandingdebt_522A':Amount of outstanding debt on the client's previous application.It can be divided into two categories based on whether it is 0 or not.\n\n- 'pmtnum_8L': Number of payments made for the previous application.\n\n- 'postype_4733339M': Type of point of sale.Here,we can consider one-hot encoding.\n\n- 'profession_152M':Profession of the client during their previous loan application.Here, 'a55475b1' accounts for the majority.\n\n- 'rejectreason_755M':Reason for previous application rejection.Here,'a55475b1' accounts for the majority.\n\n- 'rejectreasonclient_4145042M':Reason for rejection of the client's previous application.Here,'a55475b1' accounts for the majority.\n\n- 'revolvingaccount_394A':Revolving account that was present in the applicant's previous application.This feeling is not very useful.\n\n- 'status_219L':Previous application status.Here,we can consider one-hot encoding.\n\n- 'tenor_203L':Number of instalments in the previous application.Here, ordinary statistical features can be done.","metadata":{}},{"cell_type":"markdown","source":"## applprev_2.csv ","metadata":{}},{"cell_type":"code","source":"applprev2=pd.read_csv(\"/kaggle/input/home-credit-credit-risk-model-stability/csv_files/train/train_applprev_2.csv\")\napplprev2","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### The official explanation is as follows: ","metadata":{}},{"cell_type":"code","source":"tmp=feature[feature['Variable'].isin(applprev2.columns)]\ntmp","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"\nunique=['PRIMARY_MOBILE', 'EMPLOYMENT_PHONE', 'PHONE', 'PRIMARY_EMAIL',\n       'HOME_PHONE', 'SECONDARY_MOBILE', 'ALTERNATIVE_PHONE', nan,\n       'WHATSAPP', 'SKYPE']\n['主要手机'，'工作电话'，'电话'，'私人邮箱'，“家庭电话”、“次要手机”、“备用电话”、nan、“什么应用”，“网络电话”]\nunique=[nan, 'CANCELLED', 'INACTIVE', 'ACTIVE', 'BLOCKED', 'RENEWED','UNCONFIRMED']\n“已取消”、“未激活”、“活动”、“已阻止”、“续订”、“尚未确认”\n\"\"\"\n\nfor v in tmp['Variable'].values:\n    print(f\"{v}.unique():{applprev2[v].unique()}\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Variable interpretation:\n\n- 'cacccardblochreas_147M': Card blocking reason.Credit card blockage is usually caused by the bank freezing your credit card, which may be due to your failure to repay the money or improper use of the credit card. There is no explanation for the string here online, and it may be due to an internal string. Is the missing string possibly due to not being frozen or not having a credit card?\n\n- 'conts_type_509L':This is the contact information left by someone when they applied for a credit card before.Missing values may be due to not leaving contact information,This is a new state.\n\n- 'credacc_cards_status_52L':The status of the previous credit card.The missing value here may be that there is no card. However,we cannot determine what the missing value is, so we can consider filling it in with 'UNCONFIRMED'.\n\n- 'case_id','num_group1','num_group2':I won't explain here.\n\n### My suggestion:\n\n- Through this table, we can classify people into several categories: those without credit cards(cacccardblochreas_147M==nan and conts_type_509L==nan), those with unfrozen credit cards(cacccardblochreas_147M==nan and conts_type_509L!=nan), and those with frozen credit cards(cacccardblochreas_147M!=nan).\n\n- Since a 'case_id' has many records, we only need the latest contact information for each 'case_id'.\n\n- For credit card freezing, it is necessary to count how many times it has been frozen and what are the reasons for each.\n\n- The status of the previous credit card can be considered using unique hot codes or filled in with the mean of the target .","metadata":{}},{"cell_type":"markdown","source":"#### The content insights of the remaining files will be discussed in the future.","metadata":{}}]}