{
  "id": 507946,
  "title": "Public 8th/Private 253 solution ",
  "url": "/competitions/home-credit-credit-risk-model-stability/writeups/evgeniia-grigoreva-public-8th-private-253-solution",
  "author_name": "",
  "post_date": "2024-05-28T10:01:00.280Z",
  "votes": 60,
  "comment_count": 13,
  "views": 0,
  "content": "<p><a href=\"https://github.com/evgeniavolkova/kagglehomecredit2024\" target=\"_blank\">Link to github repo with the code.</a></p>\n<p>This is the write-up for my part of the solution. My teammate Vitaly will add his solution later.</p>\n<p>What a great shakup! And it seems that the key to achieving a good score on private LB was using a correct approach to hack the metric. </p>\n<p>I restored date information using the \"refreshdate_3813885D\" column from the credit_bureau_a1* files. The minimum value per case_id is almost always 3/1/2019. The difference between \"refreshdate_3813885D\" and \"date_decision\" has a near-perfect correlation with \"date_decision\". Since date differences are preserved in the test set, I restored the original date by subtracting this difference from 3/1/2019. See <a href=\"https://www.kaggle.com/code/eivolkova/how-to-restore-dates/notebook\" target=\"_blank\">notebook</a> for details.</p>\n<p>The percentage of correctly restored dates in the training dataset was 87%, with 3% errors and 10% missing values.</p>\n<p>This method helped me achieve a maximum score of 0.655 on the public LB. However, the publicly shared method worked better on LB, giving a +0.01 boost. On the private LB the original method worked better, giving a score of 0.560-0.617, while the public method gave only 0.510-0.520. My best private without hack was around 0.520, so the public version of the hack didn't help.</p>\n<p>Upd: for some reason, I was sure that the private LB included the public LB, so I thought any hack would work. It seems that's not the case, which explains why restoring the dates and reducing AUC for the first year worked, while the public hack didn’t.</p>\n<p>A few more words on the metric: hackability was a significant issue, but not the only one. The linear time trend penalty included in the metric can be substantial when there's no actual predictable trend in the ginis. I discussed this issue in more detail <a href=\"https://www.kaggle.com/code/eivolkova/stability-metric-issue-probabilistic-approach/notebook\" target=\"_blank\">here</a>. I strongly believe that the downward slope observed in the test set is related to changing market conditions rather than the deterioration of the model's prediction quality. These changes in market conditions create a stochastic trend, which is almost always observed in non-differentiated financial time series and is not predictable. Therefore, extrapolating this downward slope is not valid.</p>\n<h3>CV strategy</h3>\n<p>In this competition, it was incredibly difficult to match CV and LB scores. I tried various approaches: simple cross-validation, time series CV, time series CV with a gap, and using only hold-out samples from weeks 80, 65, 50… The problem was that nothing correlated with the LB, and there was no negative slope to estimate the model's stability. </p>\n<p>The low correlation with LB could be due to the fact that the post-COVID period was different from the periods before. My hypothesis is that this might be because, at some point, government support started to reduce, which changed people's cash flows. Additionally, in 2022, the economy was hit by another shock. There could also have been distortions in the data due to changes in reporting because of the CARES Act: if a person was late in payments, the lender was required to maintain their delinquency status during the accommodation period but could not report them as more delinquent.</p>\n<p>My final approach was to split the data into five folds sequentially based on weeks, and drop slope term from the metric as it didn't have any meaning withing this CV scheme.</p>\n<h4>Finding a slope</h4>\n<p>I also tried a fancy approach to replicate the downward slope. I set aside the last 40 weeks (50-91) as a holdout set and trained two models: one on the first 25 weeks (0-25) and the second on the next 25 weeks (25-50). This was to estimate how much the model performance drops when trained on the most recent data compared to older data. Then I concatenated predictions from the second model with predictions from the first model and calculated the ginis stability metric, which now had a negative slope penalty. However, this approach also didn't correlate with LB scores, so I only considered it for feature selection purposes for a few models.</p>\n<h3>Feature engineering</h3>\n<ul>\n<li>First, I reviewed the data and all four hundred features provided, attempting to create meaningful features based on them (ratios, differences, etc.) and aggregations. I won't list them all here, but they can be found in my <a href=\"https://github.com/evgeniavolkova/kagglehomecredit2024\" target=\"_blank\">repo</a>.</li>\n<li><strong>Person Data</strong>: I noticed that when num_group1=0, the data pertains to the applicant, while other values correspond to related persons. Thus, I extracted data with num_group1=0 and added it to the feature set, applying aggregations for the rest.</li>\n<li><strong>Credit Bureau Data</strong>: I noticed that there are duplicate columns for active and closed contracts, so I merged each pair into one to build aggregations based on all contracts, retaining active and closed contract aggregations as well.</li>\n<li><strong>Time Aggregations</strong>: For credit bureau data, I created aggregations over different time intervals. For contracts data, I used contracts ending in the last 3, 5, or 7 years. For installment data, I used 1 month, 6 months, and 1, 2, or 3 years. This helped to improve the CV score significantly. I haven't tried aggregations based on the start of the credit, maybe it could yield even better results.</li>\n<li><strong>Categorical Features Aggregations</strong>: I used mode, unique values count, and a measure of diversity (number of unique values divided by the total number per each case_id).</li>\n<li><strong>Same Data from Different Source</strong>: I concatenated birth date, employment length, taxes, and other columns that were present in the dataset more than once. This improved CV scores; however, removing the original columns reduced CV scores. A better strategy might be to calculate the mode from all sources.</li>\n<li><strong>Credit Bureau Sources Concatenation</strong>: As there are two sources for previous contracts data (a and b), I manually concatenated them. However, this didn't significantly improve results, so I kept this only in one model.</li>\n<li><strong>DPD features</strong>: For previous applications and credit bureau data I created aggregations only for those rows where DPD is different from zero.</li>\n</ul>\n<h2>Feature selection</h2>\n<p>For feature selection, I used two approaches: forward feature selection and null feature importances (borrowed from <a href=\"https://www.kaggle.com/code/ogrellier/feature-selection-with-null-importances\" target=\"_blank\">previous competition notebook</a>).</p>\n<h2>Ensembling</h2>\n<p>When I started ensembling models, the correlation between CV and LB improved. It wasn't perfect, but ensembling finally allowed me to achieve LB scores higher than 0.571. I realized that training a few different tree-based models like LightGBM, CatBoost, and XGBoost wasn't enough. Initially, I used around 300 features selected by forward feature selection and added random subsets of features with different parameters for LightGBM and various seeds. Then I understood that to achieve even greater diversity, I needed to repeat the feature engineering process from scratch, create a new dataset, train a model on it, and add it to the existing ensemble. This resulted in having six different data processors, significantly enhancing the diversity of the solutions. In the end, I used an ensemble which gave me a good CV score and the highest LB score I could achieve. </p>\n<h2>Models</h2>\n<p>I used only LightGBM and CatBoost. I tried to build a few simple DL models, however, they didn't add anything to my ensemble, so I dropped them.</p>\n<h2>What didn't work</h2>\n<ul>\n<li><strong>Submodels</strong>: I attempted to follow past Home Credit competition winners' approach by creating submodels for each dataset with depth &gt;0. For each dataset I added 'target' column from train_base, trained models on CV, and used different aggregations of OOF predictions as meta-features (see this <a href=\"https://www.kaggle.com/competitions/home-credit-default-risk/discussion/64596\" target=\"_blank\">discussion</a>). This only worked for credit bureau data on CV, and it significantly increased time and RAM requirements, so I abandoned this approach.</li>\n<li><strong>Scaling</strong>: I was concerned about the non-stationarity of many features, primarily due to inflation for amount columns. Scaling by inflation measures worsened CV. Scaling all amount columns by income didn't work as well, even when I used scaled features in addition to original ones. This might be due to unreliable data on income: I noticed low correlation between income data from person1 and static0 sources (see this <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/491927\" target=\"_blank\">discussion</a>). This might be due to data provider post-processing or applicant mistakes.</li>\n<li><strong>Other approaches</strong> that didn't work: average target for N closest neighbors; pseudo-labeling; augmenting the training dataset with data from previous applications; creating dummies for categorical features from depth&gt;0 files and summing up; weightning observations with higher weight on recent ones; excluding features based on adversarial validation; weighted average of depth&gt;0 columns based on time; creating seperate models for the case_ids with credit history and without it.</li>\n</ul>",
  "messages": [
    {
      "id": "2840055",
      "postDate": "05/28/2024 00:15:54",
      "content": "<p><a href=\"https://github.com/evgeniavolkova/kagglehomecredit2024\" target=\"_blank\">Link to github repo with the code.</a></p>\n<p>This is the write-up for my part of the solution. My teammate Vitaly will add his solution later.</p>\n<p>What a great shakup! And it seems that the key to achieving a good score on private LB was using a correct approach to hack the metric. </p>\n<p>I restored date information using the \"refreshdate_3813885D\" column from the credit_bureau_a1* files. The minimum value per case_id is almost always 3/1/2019. The difference between \"refreshdate_3813885D\" and \"date_decision\" has a near-perfect correlation with \"date_decision\". Since date differences are preserved in the test set, I restored the original date by subtracting this difference from 3/1/2019. See <a href=\"https://www.kaggle.com/code/eivolkova/how-to-restore-dates/notebook\" target=\"_blank\">notebook</a> for details.</p>\n<p>The percentage of correctly restored dates in the training dataset was 87%, with 3% errors and 10% missing values.</p>\n<p>This method helped me achieve a maximum score of 0.655 on the public LB. However, the publicly shared method worked better on LB, giving a +0.01 boost. On the private LB the original method worked better, giving a score of 0.560-0.617, while the public method gave only 0.510-0.520. My best private without hack was around 0.520, so the public version of the hack didn't help.</p>\n<p>Upd: for some reason, I was sure that the private LB included the public LB, so I thought any hack would work. It seems that's not the case, which explains why restoring the dates and reducing AUC for the first year worked, while the public hack didn’t.</p>\n<p>A few more words on the metric: hackability was a significant issue, but not the only one. The linear time trend penalty included in the metric can be substantial when there's no actual predictable trend in the ginis. I discussed this issue in more detail <a href=\"https://www.kaggle.com/code/eivolkova/stability-metric-issue-probabilistic-approach/notebook\" target=\"_blank\">here</a>. I strongly believe that the downward slope observed in the test set is related to changing market conditions rather than the deterioration of the model's prediction quality. These changes in market conditions create a stochastic trend, which is almost always observed in non-differentiated financial time series and is not predictable. Therefore, extrapolating this downward slope is not valid.</p>\n<h3>CV strategy</h3>\n<p>In this competition, it was incredibly difficult to match CV and LB scores. I tried various approaches: simple cross-validation, time series CV, time series CV with a gap, and using only hold-out samples from weeks 80, 65, 50… The problem was that nothing correlated with the LB, and there was no negative slope to estimate the model's stability. </p>\n<p>The low correlation with LB could be due to the fact that the post-COVID period was different from the periods before. My hypothesis is that this might be because, at some point, government support started to reduce, which changed people's cash flows. Additionally, in 2022, the economy was hit by another shock. There could also have been distortions in the data due to changes in reporting because of the CARES Act: if a person was late in payments, the lender was required to maintain their delinquency status during the accommodation period but could not report them as more delinquent.</p>\n<p>My final approach was to split the data into five folds sequentially based on weeks, and drop slope term from the metric as it didn't have any meaning withing this CV scheme.</p>\n<h4>Finding a slope</h4>\n<p>I also tried a fancy approach to replicate the downward slope. I set aside the last 40 weeks (50-91) as a holdout set and trained two models: one on the first 25 weeks (0-25) and the second on the next 25 weeks (25-50). This was to estimate how much the model performance drops when trained on the most recent data compared to older data. Then I concatenated predictions from the second model with predictions from the first model and calculated the ginis stability metric, which now had a negative slope penalty. However, this approach also didn't correlate with LB scores, so I only considered it for feature selection purposes for a few models.</p>\n<h3>Feature engineering</h3>\n<ul>\n<li>First, I reviewed the data and all four hundred features provided, attempting to create meaningful features based on them (ratios, differences, etc.) and aggregations. I won't list them all here, but they can be found in my <a href=\"https://github.com/evgeniavolkova/kagglehomecredit2024\" target=\"_blank\">repo</a>.</li>\n<li><strong>Person Data</strong>: I noticed that when num_group1=0, the data pertains to the applicant, while other values correspond to related persons. Thus, I extracted data with num_group1=0 and added it to the feature set, applying aggregations for the rest.</li>\n<li><strong>Credit Bureau Data</strong>: I noticed that there are duplicate columns for active and closed contracts, so I merged each pair into one to build aggregations based on all contracts, retaining active and closed contract aggregations as well.</li>\n<li><strong>Time Aggregations</strong>: For credit bureau data, I created aggregations over different time intervals. For contracts data, I used contracts ending in the last 3, 5, or 7 years. For installment data, I used 1 month, 6 months, and 1, 2, or 3 years. This helped to improve the CV score significantly. I haven't tried aggregations based on the start of the credit, maybe it could yield even better results.</li>\n<li><strong>Categorical Features Aggregations</strong>: I used mode, unique values count, and a measure of diversity (number of unique values divided by the total number per each case_id).</li>\n<li><strong>Same Data from Different Source</strong>: I concatenated birth date, employment length, taxes, and other columns that were present in the dataset more than once. This improved CV scores; however, removing the original columns reduced CV scores. A better strategy might be to calculate the mode from all sources.</li>\n<li><strong>Credit Bureau Sources Concatenation</strong>: As there are two sources for previous contracts data (a and b), I manually concatenated them. However, this didn't significantly improve results, so I kept this only in one model.</li>\n<li><strong>DPD features</strong>: For previous applications and credit bureau data I created aggregations only for those rows where DPD is different from zero.</li>\n</ul>\n<h2>Feature selection</h2>\n<p>For feature selection, I used two approaches: forward feature selection and null feature importances (borrowed from <a href=\"https://www.kaggle.com/code/ogrellier/feature-selection-with-null-importances\" target=\"_blank\">previous competition notebook</a>).</p>\n<h2>Ensembling</h2>\n<p>When I started ensembling models, the correlation between CV and LB improved. It wasn't perfect, but ensembling finally allowed me to achieve LB scores higher than 0.571. I realized that training a few different tree-based models like LightGBM, CatBoost, and XGBoost wasn't enough. Initially, I used around 300 features selected by forward feature selection and added random subsets of features with different parameters for LightGBM and various seeds. Then I understood that to achieve even greater diversity, I needed to repeat the feature engineering process from scratch, create a new dataset, train a model on it, and add it to the existing ensemble. This resulted in having six different data processors, significantly enhancing the diversity of the solutions. In the end, I used an ensemble which gave me a good CV score and the highest LB score I could achieve. </p>\n<h2>Models</h2>\n<p>I used only LightGBM and CatBoost. I tried to build a few simple DL models, however, they didn't add anything to my ensemble, so I dropped them.</p>\n<h2>What didn't work</h2>\n<ul>\n<li><strong>Submodels</strong>: I attempted to follow past Home Credit competition winners' approach by creating submodels for each dataset with depth &gt;0. For each dataset I added 'target' column from train_base, trained models on CV, and used different aggregations of OOF predictions as meta-features (see this <a href=\"https://www.kaggle.com/competitions/home-credit-default-risk/discussion/64596\" target=\"_blank\">discussion</a>). This only worked for credit bureau data on CV, and it significantly increased time and RAM requirements, so I abandoned this approach.</li>\n<li><strong>Scaling</strong>: I was concerned about the non-stationarity of many features, primarily due to inflation for amount columns. Scaling by inflation measures worsened CV. Scaling all amount columns by income didn't work as well, even when I used scaled features in addition to original ones. This might be due to unreliable data on income: I noticed low correlation between income data from person1 and static0 sources (see this <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/491927\" target=\"_blank\">discussion</a>). This might be due to data provider post-processing or applicant mistakes.</li>\n<li><strong>Other approaches</strong> that didn't work: average target for N closest neighbors; pseudo-labeling; augmenting the training dataset with data from previous applications; creating dummies for categorical features from depth&gt;0 files and summing up; weightning observations with higher weight on recent ones; excluding features based on adversarial validation; weighted average of depth&gt;0 columns based on time; creating seperate models for the case_ids with credit history and without it.</li>\n</ul>",
      "rawMarkdown": "[Link to github repo with the code.](https://github.com/evgeniavolkova/kagglehomecredit2024)\n\nThis is the write-up for my part of the solution. My teammate Vitaly will add his solution later.\n\nWhat a great shakup! And it seems that the key to achieving a good score on private LB was using a correct approach to hack the metric. \n\nI restored date information using the \"refreshdate_3813885D\" column from the credit_bureau_a1* files. The minimum value per case_id is almost always 3/1/2019. The difference between \"refreshdate_3813885D\" and \"date_decision\" has a near-perfect correlation with \"date_decision\". Since date differences are preserved in the test set, I restored the original date by subtracting this difference from 3/1/2019. See [notebook](https://www.kaggle.com/code/eivolkova/how-to-restore-dates/notebook) for details.\n\nThe percentage of correctly restored dates in the training dataset was 87%, with 3% errors and 10% missing values.\n\nThis method helped me achieve a maximum score of 0.655 on the public LB. However, the publicly shared method worked better on LB, giving a +0.01 boost. On the private LB the original method worked better, giving a score of 0.560-0.617, while the public method gave only 0.510-0.520. My best private without hack was around 0.520, so the public version of the hack didn't help.\n\nUpd: for some reason, I was sure that the private LB included the public LB, so I thought any hack would work. It seems that's not the case, which explains why restoring the dates and reducing AUC for the first year worked, while the public hack didn’t.\n\nA few more words on the metric: hackability was a significant issue, but not the only one. The linear time trend penalty included in the metric can be substantial when there's no actual predictable trend in the ginis. I discussed this issue in more detail [here](https://www.kaggle.com/code/eivolkova/stability-metric-issue-probabilistic-approach/notebook). I strongly believe that the downward slope observed in the test set is related to changing market conditions rather than the deterioration of the model's prediction quality. These changes in market conditions create a stochastic trend, which is almost always observed in non-differentiated financial time series and is not predictable. Therefore, extrapolating this downward slope is not valid.\n\n### CV strategy\nIn this competition, it was incredibly difficult to match CV and LB scores. I tried various approaches: simple cross-validation, time series CV, time series CV with a gap, and using only hold-out samples from weeks 80, 65, 50... The problem was that nothing correlated with the LB, and there was no negative slope to estimate the model's stability. \n\nThe low correlation with LB could be due to the fact that the post-COVID period was different from the periods before. My hypothesis is that this might be because, at some point, government support started to reduce, which changed people's cash flows. Additionally, in 2022, the economy was hit by another shock. There could also have been distortions in the data due to changes in reporting because of the CARES Act: if a person was late in payments, the lender was required to maintain their delinquency status during the accommodation period but could not report them as more delinquent.\n\nMy final approach was to split the data into five folds sequentially based on weeks, and drop slope term from the metric as it didn't have any meaning withing this CV scheme.\n\n#### Finding a slope\nI also tried a fancy approach to replicate the downward slope. I set aside the last 40 weeks (50-91) as a holdout set and trained two models: one on the first 25 weeks (0-25) and the second on the next 25 weeks (25-50). This was to estimate how much the model performance drops when trained on the most recent data compared to older data. Then I concatenated predictions from the second model with predictions from the first model and calculated the ginis stability metric, which now had a negative slope penalty. However, this approach also didn't correlate with LB scores, so I only considered it for feature selection purposes for a few models.\n\n### Feature engineering\n- First, I reviewed the data and all four hundred features provided, attempting to create meaningful features based on them (ratios, differences, etc.) and aggregations. I won't list them all here, but they can be found in my [repo](https://github.com/evgeniavolkova/kagglehomecredit2024).\n- **Person Data**: I noticed that when num_group1=0, the data pertains to the applicant, while other values correspond to related persons. Thus, I extracted data with num_group1=0 and added it to the feature set, applying aggregations for the rest.\n- **Credit Bureau Data**: I noticed that there are duplicate columns for active and closed contracts, so I merged each pair into one to build aggregations based on all contracts, retaining active and closed contract aggregations as well.\n- **Time Aggregations**: For credit bureau data, I created aggregations over different time intervals. For contracts data, I used contracts ending in the last 3, 5, or 7 years. For installment data, I used 1 month, 6 months, and 1, 2, or 3 years. This helped to improve the CV score significantly. I haven't tried aggregations based on the start of the credit, maybe it could yield even better results.\n- **Categorical Features Aggregations**: I used mode, unique values count, and a measure of diversity (number of unique values divided by the total number per each case_id).\n- **Same Data from Different Source**: I concatenated birth date, employment length, taxes, and other columns that were present in the dataset more than once. This improved CV scores; however, removing the original columns reduced CV scores. A better strategy might be to calculate the mode from all sources.\n- **Credit Bureau Sources Concatenation**: As there are two sources for previous contracts data (a and b), I manually concatenated them. However, this didn't significantly improve results, so I kept this only in one model.\n- **DPD features**: For previous applications and credit bureau data I created aggregations only for those rows where DPD is different from zero.\n\n## Feature selection\nFor feature selection, I used two approaches: forward feature selection and null feature importances (borrowed from [previous competition notebook](https://www.kaggle.com/code/ogrellier/feature-selection-with-null-importances)).\n\n## Ensembling\nWhen I started ensembling models, the correlation between CV and LB improved. It wasn't perfect, but ensembling finally allowed me to achieve LB scores higher than 0.571. I realized that training a few different tree-based models like LightGBM, CatBoost, and XGBoost wasn't enough. Initially, I used around 300 features selected by forward feature selection and added random subsets of features with different parameters for LightGBM and various seeds. Then I understood that to achieve even greater diversity, I needed to repeat the feature engineering process from scratch, create a new dataset, train a model on it, and add it to the existing ensemble. This resulted in having six different data processors, significantly enhancing the diversity of the solutions. In the end, I used an ensemble which gave me a good CV score and the highest LB score I could achieve. \n\n## Models\nI used only LightGBM and CatBoost. I tried to build a few simple DL models, however, they didn't add anything to my ensemble, so I dropped them.\n\n## What didn't work\n- **Submodels**: I attempted to follow past Home Credit competition winners' approach by creating submodels for each dataset with depth >0. For each dataset I added 'target' column from train_base, trained models on CV, and used different aggregations of OOF predictions as meta-features (see this [discussion](https://www.kaggle.com/competitions/home-credit-default-risk/discussion/64596)). This only worked for credit bureau data on CV, and it significantly increased time and RAM requirements, so I abandoned this approach.\n- **Scaling**: I was concerned about the non-stationarity of many features, primarily due to inflation for amount columns. Scaling by inflation measures worsened CV. Scaling all amount columns by income didn't work as well, even when I used scaled features in addition to original ones. This might be due to unreliable data on income: I noticed low correlation between income data from person1 and static0 sources (see this [discussion](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/491927)). This might be due to data provider post-processing or applicant mistakes.\n- **Other approaches** that didn't work: average target for N closest neighbors; pseudo-labeling; augmenting the training dataset with data from previous applications; creating dummies for categorical features from depth>0 files and summing up; weightning observations with higher weight on recent ones; excluding features based on adversarial validation; weighted average of depth>0 columns based on time; creating seperate models for the case_ids with credit history and without it.",
      "votes": null
    },
    {
      "id": "2840073",
      "postDate": "05/28/2024 00:39:54",
      "content": "<p>Your sharing will be a great momentum to help DS like me improve our skills. Thanks for sharing.</p>\n<p>May I ask how metric hacking impact to your both public LB and private LB? Did you spend time to optimize metric hacking?</p>",
      "rawMarkdown": "Your sharing will be a great momentum to help DS like me improve our skills. Thanks for sharing.\n\nMay I ask how metric hacking impact to your both public LB and private LB? Did you spend time to optimize metric hacking?",
      "votes": null
    },
    {
      "id": "2840178",
      "postDate": "05/28/2024 02:19:58",
      "content": "<p>Thank you for your detailed solution!</p>",
      "rawMarkdown": "Thank you for your detailed solution!",
      "votes": null
    },
    {
      "id": "2840261",
      "postDate": "05/28/2024 03:59:24",
      "content": "<p>Very good analysis, I learned a lot from your post. Thank you for sharing your insights!<br>\nMy team was busted by the shake though…<br>\nJust curious, do you believe this competition helped us to improve our DS skills?  </p>",
      "rawMarkdown": "Very good analysis, I learned a lot from your post. Thank you for sharing your insights!\nMy team was busted by the shake though...\nJust curious, do you believe this competition helped us to improve our DS skills?",
      "votes": null
    },
    {
      "id": "2840702",
      "postDate": "05/28/2024 08:09:26",
      "content": "<p>Nice work! Comment from bronze medal winner</p>",
      "rawMarkdown": "Nice work! Comment from bronze medal winner",
      "votes": null
    },
    {
      "id": "2840829",
      "postDate": "05/28/2024 09:15:00",
      "content": "<p>Thank you for idea.</p>",
      "rawMarkdown": "Thank you for idea.",
      "votes": null
    },
    {
      "id": "2840885",
      "postDate": "05/28/2024 09:39:05",
      "content": "<p>I believe it helped me because I was in the competition from the start and tried to build good models before everyone started to use the hack.</p>\n<p>But yes, the hack spoiled everything. The key to winning wasn’t even about finding a way to hack the metric but rather probing the leaderboard as much as possible to gather information about the public/private split. We could call it a life lesson, but definitely not a data science lesson. :)</p>",
      "rawMarkdown": "I believe it helped me because I was in the competition from the start and tried to build good models before everyone started to use the hack.\n\nBut yes, the hack spoiled everything. The key to winning wasn’t even about finding a way to hack the metric but rather probing the leaderboard as much as possible to gather information about the public/private split. We could call it a life lesson, but definitely not a data science lesson. :)",
      "votes": null
    },
    {
      "id": "2840892",
      "postDate": "05/28/2024 09:42:55",
      "content": "<p>Public hack didn't help to increase the score (maybe it has even reduced a bit) on the private LB, but it increased the score on the public LB by ~0.6.<br>\nRecovering dates and reducing the accuracy of predictions of the first ~12 months helped to achieve 0.6-0.617 private and ~0.650 public. </p>",
      "rawMarkdown": "Public hack didn't help to increase the score (maybe it has even reduced a bit) on the private LB, but it increased the score on the public LB by ~0.6.\nRecovering dates and reducing the accuracy of predictions of the first ~12 months helped to achieve 0.6-0.617 private and ~0.650 public.",
      "votes": null
    },
    {
      "id": "2841831",
      "postDate": "05/28/2024 17:50:54",
      "content": "<p>So you had potential submissions for 1st place, but you and your teammate chose others?</p>",
      "rawMarkdown": "So you had potential submissions for 1st place, but you and your teammate chose others?",
      "votes": null
    },
    {
      "id": "2841863",
      "postDate": "05/28/2024 18:07:01",
      "content": "<p>Yeah, as many others, I think. I thought that public is a part of private, which means that any hack should work on private, if it works on public. From probling it looked like the split is more or less random. </p>",
      "rawMarkdown": "Yeah, as many others, I think. I thought that public is a part of private, which means that any hack should work on private, if it works on public. From probling it looked like the split is more or less random.",
      "votes": null
    },
    {
      "id": "2841991",
      "postDate": "05/28/2024 19:44:20",
      "content": "<p>Nice work!!</p>",
      "rawMarkdown": "Nice work!!",
      "votes": null
    },
    {
      "id": "2842569",
      "postDate": "05/29/2024 06:49:18",
      "content": "<p>Your github repo is very nice!</p>",
      "rawMarkdown": "Your github repo is very nice!",
      "votes": null
    },
    {
      "id": "2842716",
      "postDate": "05/29/2024 08:18:19",
      "content": "<p>Nice Work!! The insights you provied are really valuable specially the part about the models that didnt work helped me understand the approach better.</p>",
      "rawMarkdown": "Nice Work!! The insights you provied are really valuable specially the part about the models that didnt work helped me understand the approach better.",
      "votes": null
    },
    {
      "id": "2920274",
      "postDate": "07/13/2024 14:50:33",
      "content": "<p>totally love it.</p>",
      "rawMarkdown": "totally love it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2840073,
      "author_name": "truythu169",
      "author_url": "",
      "post_date": "05/28/2024 00:39:54",
      "content": "<p>Your sharing will be a great momentum to help DS like me improve our skills. Thanks for sharing.</p>\n<p>May I ask how metric hacking impact to your both public LB and private LB? Did you spend time to optimize metric hacking?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2840892,
          "author_name": "eivolkova",
          "author_url": "",
          "post_date": "05/28/2024 09:42:55",
          "content": "<p>Public hack didn't help to increase the score (maybe it has even reduced a bit) on the private LB, but it increased the score on the public LB by ~0.6.<br>\nRecovering dates and reducing the accuracy of predictions of the first ~12 months helped to achieve 0.6-0.617 private and ~0.650 public. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2841831,
              "author_name": "andreynesterov",
              "author_url": "",
              "post_date": "05/28/2024 17:50:54",
              "content": "<p>So you had potential submissions for 1st place, but you and your teammate chose others?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2841863,
                  "author_name": "eivolkova",
                  "author_url": "",
                  "post_date": "05/28/2024 18:07:01",
                  "content": "<p>Yeah, as many others, I think. I thought that public is a part of private, which means that any hack should work on private, if it works on public. From probling it looked like the split is more or less random. </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2840178,
      "author_name": "liwenqiang1218",
      "author_url": "",
      "post_date": "05/28/2024 02:19:58",
      "content": "<p>Thank you for your detailed solution!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2840261,
      "author_name": "ttkuma",
      "author_url": "",
      "post_date": "05/28/2024 03:59:24",
      "content": "<p>Very good analysis, I learned a lot from your post. Thank you for sharing your insights!<br>\nMy team was busted by the shake though…<br>\nJust curious, do you believe this competition helped us to improve our DS skills?  </p>",
      "votes": null,
      "replies": [
        {
          "id": 2840885,
          "author_name": "eivolkova",
          "author_url": "",
          "post_date": "05/28/2024 09:39:05",
          "content": "<p>I believe it helped me because I was in the competition from the start and tried to build good models before everyone started to use the hack.</p>\n<p>But yes, the hack spoiled everything. The key to winning wasn’t even about finding a way to hack the metric but rather probing the leaderboard as much as possible to gather information about the public/private split. We could call it a life lesson, but definitely not a data science lesson. :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2840702,
      "author_name": "kaixuanchen77",
      "author_url": "",
      "post_date": "05/28/2024 08:09:26",
      "content": "<p>Nice work! Comment from bronze medal winner</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2840829,
      "author_name": "pateljinalvinodbhai",
      "author_url": "",
      "post_date": "05/28/2024 09:15:00",
      "content": "<p>Thank you for idea.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2841991,
      "author_name": "savitpawar",
      "author_url": "",
      "post_date": "05/28/2024 19:44:20",
      "content": "<p>Nice work!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2842569,
      "author_name": "leewook",
      "author_url": "",
      "post_date": "05/29/2024 06:49:18",
      "content": "<p>Your github repo is very nice!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2842716,
      "author_name": "adityakalsi",
      "author_url": "",
      "post_date": "05/29/2024 08:18:19",
      "content": "<p>Nice Work!! The insights you provied are really valuable specially the part about the models that didnt work helped me understand the approach better.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2920274,
      "author_name": "ismailovelvinknown",
      "author_url": "",
      "post_date": "07/13/2024 14:50:33",
      "content": "<p>totally love it.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2840055": "[Link to github repo with the code.](https://github.com/evgeniavolkova/kagglehomecredit2024)\n\nThis is the write-up for my part of the solution. My teammate Vitaly will add his solution later.\n\nWhat a great shakup! And it seems that the key to achieving a good score on private LB was using a correct approach to hack the metric. \n\nI restored date information using the \"refreshdate_3813885D\" column from the credit_bureau_a1* files. The minimum value per case_id is almost always 3/1/2019. The difference between \"refreshdate_3813885D\" and \"date_decision\" has a near-perfect correlation with \"date_decision\". Since date differences are preserved in the test set, I restored the original date by subtracting this difference from 3/1/2019. See [notebook](https://www.kaggle.com/code/eivolkova/how-to-restore-dates/notebook) for details.\n\nThe percentage of correctly restored dates in the training dataset was 87%, with 3% errors and 10% missing values.\n\nThis method helped me achieve a maximum score of 0.655 on the public LB. However, the publicly shared method worked better on LB, giving a +0.01 boost. On the private LB the original method worked better, giving a score of 0.560-0.617, while the public method gave only 0.510-0.520. My best private without hack was around 0.520, so the public version of the hack didn't help.\n\nUpd: for some reason, I was sure that the private LB included the public LB, so I thought any hack would work. It seems that's not the case, which explains why restoring the dates and reducing AUC for the first year worked, while the public hack didn’t.\n\nA few more words on the metric: hackability was a significant issue, but not the only one. The linear time trend penalty included in the metric can be substantial when there's no actual predictable trend in the ginis. I discussed this issue in more detail [here](https://www.kaggle.com/code/eivolkova/stability-metric-issue-probabilistic-approach/notebook). I strongly believe that the downward slope observed in the test set is related to changing market conditions rather than the deterioration of the model's prediction quality. These changes in market conditions create a stochastic trend, which is almost always observed in non-differentiated financial time series and is not predictable. Therefore, extrapolating this downward slope is not valid.\n\n### CV strategy\nIn this competition, it was incredibly difficult to match CV and LB scores. I tried various approaches: simple cross-validation, time series CV, time series CV with a gap, and using only hold-out samples from weeks 80, 65, 50... The problem was that nothing correlated with the LB, and there was no negative slope to estimate the model's stability. \n\nThe low correlation with LB could be due to the fact that the post-COVID period was different from the periods before. My hypothesis is that this might be because, at some point, government support started to reduce, which changed people's cash flows. Additionally, in 2022, the economy was hit by another shock. There could also have been distortions in the data due to changes in reporting because of the CARES Act: if a person was late in payments, the lender was required to maintain their delinquency status during the accommodation period but could not report them as more delinquent.\n\nMy final approach was to split the data into five folds sequentially based on weeks, and drop slope term from the metric as it didn't have any meaning withing this CV scheme.\n\n#### Finding a slope\nI also tried a fancy approach to replicate the downward slope. I set aside the last 40 weeks (50-91) as a holdout set and trained two models: one on the first 25 weeks (0-25) and the second on the next 25 weeks (25-50). This was to estimate how much the model performance drops when trained on the most recent data compared to older data. Then I concatenated predictions from the second model with predictions from the first model and calculated the ginis stability metric, which now had a negative slope penalty. However, this approach also didn't correlate with LB scores, so I only considered it for feature selection purposes for a few models.\n\n### Feature engineering\n- First, I reviewed the data and all four hundred features provided, attempting to create meaningful features based on them (ratios, differences, etc.) and aggregations. I won't list them all here, but they can be found in my [repo](https://github.com/evgeniavolkova/kagglehomecredit2024).\n- **Person Data**: I noticed that when num_group1=0, the data pertains to the applicant, while other values correspond to related persons. Thus, I extracted data with num_group1=0 and added it to the feature set, applying aggregations for the rest.\n- **Credit Bureau Data**: I noticed that there are duplicate columns for active and closed contracts, so I merged each pair into one to build aggregations based on all contracts, retaining active and closed contract aggregations as well.\n- **Time Aggregations**: For credit bureau data, I created aggregations over different time intervals. For contracts data, I used contracts ending in the last 3, 5, or 7 years. For installment data, I used 1 month, 6 months, and 1, 2, or 3 years. This helped to improve the CV score significantly. I haven't tried aggregations based on the start of the credit, maybe it could yield even better results.\n- **Categorical Features Aggregations**: I used mode, unique values count, and a measure of diversity (number of unique values divided by the total number per each case_id).\n- **Same Data from Different Source**: I concatenated birth date, employment length, taxes, and other columns that were present in the dataset more than once. This improved CV scores; however, removing the original columns reduced CV scores. A better strategy might be to calculate the mode from all sources.\n- **Credit Bureau Sources Concatenation**: As there are two sources for previous contracts data (a and b), I manually concatenated them. However, this didn't significantly improve results, so I kept this only in one model.\n- **DPD features**: For previous applications and credit bureau data I created aggregations only for those rows where DPD is different from zero.\n\n## Feature selection\nFor feature selection, I used two approaches: forward feature selection and null feature importances (borrowed from [previous competition notebook](https://www.kaggle.com/code/ogrellier/feature-selection-with-null-importances)).\n\n## Ensembling\nWhen I started ensembling models, the correlation between CV and LB improved. It wasn't perfect, but ensembling finally allowed me to achieve LB scores higher than 0.571. I realized that training a few different tree-based models like LightGBM, CatBoost, and XGBoost wasn't enough. Initially, I used around 300 features selected by forward feature selection and added random subsets of features with different parameters for LightGBM and various seeds. Then I understood that to achieve even greater diversity, I needed to repeat the feature engineering process from scratch, create a new dataset, train a model on it, and add it to the existing ensemble. This resulted in having six different data processors, significantly enhancing the diversity of the solutions. In the end, I used an ensemble which gave me a good CV score and the highest LB score I could achieve. \n\n## Models\nI used only LightGBM and CatBoost. I tried to build a few simple DL models, however, they didn't add anything to my ensemble, so I dropped them.\n\n## What didn't work\n- **Submodels**: I attempted to follow past Home Credit competition winners' approach by creating submodels for each dataset with depth >0. For each dataset I added 'target' column from train_base, trained models on CV, and used different aggregations of OOF predictions as meta-features (see this [discussion](https://www.kaggle.com/competitions/home-credit-default-risk/discussion/64596)). This only worked for credit bureau data on CV, and it significantly increased time and RAM requirements, so I abandoned this approach.\n- **Scaling**: I was concerned about the non-stationarity of many features, primarily due to inflation for amount columns. Scaling by inflation measures worsened CV. Scaling all amount columns by income didn't work as well, even when I used scaled features in addition to original ones. This might be due to unreliable data on income: I noticed low correlation between income data from person1 and static0 sources (see this [discussion](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/491927)). This might be due to data provider post-processing or applicant mistakes.\n- **Other approaches** that didn't work: average target for N closest neighbors; pseudo-labeling; augmenting the training dataset with data from previous applications; creating dummies for categorical features from depth>0 files and summing up; weightning observations with higher weight on recent ones; excluding features based on adversarial validation; weighted average of depth>0 columns based on time; creating seperate models for the case_ids with credit history and without it.",
    "2840073": "Your sharing will be a great momentum to help DS like me improve our skills. Thanks for sharing.\n\nMay I ask how metric hacking impact to your both public LB and private LB? Did you spend time to optimize metric hacking?",
    "2840178": "Thank you for your detailed solution!",
    "2840261": "Very good analysis, I learned a lot from your post. Thank you for sharing your insights!\nMy team was busted by the shake though...\nJust curious, do you believe this competition helped us to improve our DS skills?",
    "2840702": "Nice work! Comment from bronze medal winner",
    "2840829": "Thank you for idea.",
    "2840885": "I believe it helped me because I was in the competition from the start and tried to build good models before everyone started to use the hack.\n\nBut yes, the hack spoiled everything. The key to winning wasn’t even about finding a way to hack the metric but rather probing the leaderboard as much as possible to gather information about the public/private split. We could call it a life lesson, but definitely not a data science lesson. :)",
    "2840892": "Public hack didn't help to increase the score (maybe it has even reduced a bit) on the private LB, but it increased the score on the public LB by ~0.6.\nRecovering dates and reducing the accuracy of predictions of the first ~12 months helped to achieve 0.6-0.617 private and ~0.650 public.",
    "2841831": "So you had potential submissions for 1st place, but you and your teammate chose others?",
    "2841863": "Yeah, as many others, I think. I thought that public is a part of private, which means that any hack should work on private, if it works on public. From probling it looked like the split is more or less random.",
    "2841991": "Nice work!!",
    "2842569": "Your github repo is very nice!",
    "2842716": "Nice Work!! The insights you provied are really valuable specially the part about the models that didnt work helped me understand the approach better.",
    "2920274": "totally love it."
  },
  "source": "meta"
}