{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":50160,"databundleVersionId":7921029,"sourceType":"competition"}],"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### Useful links:\n\n+ https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/overview\n+ https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/482474\n+ https://en.wikipedia.org/wiki/Credit_bureau\n+ https://www.kaggle.com/code/jetakow/home-credit-2024-starter-notebook\n+ https://stats.stackexchange.com/questions/342329/gini-and-auc-relationship\n\nBinary classification task - predict credit risk (whether given person will default).\n\nData is conveniently split into `train` and `test` subsets, indicated by prefix.\nFeatures - have several levels, single \"data point\" relates to a `case_id` (credit case, I assume).\nBase table `csv_files/train/train_base.csv` - stores `case_id`, `WEEK_NUM`, `target` and other metadata.\n\nThere is a LOT of features (`30Gb` worth of tables) - impossible to cover those in a short summary.\nFeature definitions are located in `feature_definitions.csv` (as discussed on the lecture).\n\nMetric used for ranking is a custom one, aimed on model stability: first, `gini` index is computed for each `WEEK_NUM` of the split, then OLS model is fit on these, and model is penalized for negative slope.\nResiduals of OLS models are also included in the final metric, along with averaged `gini` scores. So, competition authors aim at more robust models w.r.t. time.\n\n$$\ngini = 2 * AUC - 1 \\\\\n\\hat{y} = a * x + b \\\\\nresiduals = y - \\hat{y} \\\\\nL = mean(gini) + 0.88 * min(0, a) - 0.5 * std(residuals)\n$$\n\nStarter notebook contains some useful utilities for data loading and processing for tabular data and working example of `lgbm` model training.\n\nOne important note is that some people tried metric hacking and thus competition holders decided to hide `WEEK_NUM` and other temporal features in order to prevent such activity.\n\nSome interesting notebooks:\n\n+ [AutoML pipeline with autogluon](https://www.kaggle.com/code/takumimukaiyama/automl-addingcountencoding/notebook)\n+ [Baseline notebook](https://www.kaggle.com/code/greysky/home-credit-baseline)\n+ [Curated list of AutoML papers](https://github.com/hibayesian/awesome-automl-papers)\n+ [Preprocessing and inference for ~0.573 score](https://www.kaggle.com/code/ravi20076/homecredit-starter-inference-v2)","metadata":{}}]}