{
  "id": 508753,
  "title": "69th solution of this competition without hacking metrics",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/508753",
  "author_name": "Arthur Yu",
  "post_date": "2024-05-30T20:49:23.575000",
  "votes": 15,
  "comment_count": 0,
  "views": 0,
  "content": "<h2>Overview</h2>\n<p>Our team utilized an ensemble of two models, LightGBM and CatBoost, for training and prediction. Without employing any hacks, our team achieved a best private leaderboard score of 0.524 (which was our final best score). With the use of a date hack, our team achieved a best private leaderboard score of 0.591.</p>\n<p>Notebook link: <a href=\"https://www.kaggle.com/code/arthurroland/home-credit-lgb-cat-date-hack-solution\" target=\"_blank\">https://www.kaggle.com/code/arthurroland/home-credit-lgb-cat-date-hack-solution</a></p>\n<h2>Feature engineering</h2>\n<p>For the LightGBM model, we selected 610 features, while for the CatBoost model, we chose 535 features. </p>\n<p>Here are some feature processing methods summarized by our team that can improve model performance:</p>\n<ol>\n<li><p>Our team found that features with a missing rate between 0.9 and 0.98 contained a significant amount of useful information. By adjusting the filtering threshold from 0.9 to 0.98, we were able to significantly improve the model's performance (from an initial score of 0.585 to 0.594). </p></li>\n<li><p>We processed the date features by creating a new feature representing the number of days between the given date and January 1, 1900, which also significantly enhanced the model's performance (improving the LightGBM single model score from 0.580 to 0.585).</p></li>\n<li><p>For string-type features with more than 200 unique categories, we retained the top 20 most frequent values and set the remaining values to null. Additionally, string-type features with more than 10,000 unique categories were primarily names, which we deemed to have no value. Therefore, we directly removed these features.</p></li>\n</ol>\n<p>Aside from the methods mentioned above, our team also performed minor feature processing. However, the improvements gained were minimal and can therefore be disregarded.</p>\n<h2>Hyperparameter tuning</h2>\n<p>Our team manually set the main parameters of the model, such as depth and num_leaves. The remaining parameters were adjusted using Optuna, and we ultimately obtained the following parameter settings.</p>\n<p>The reason for manually adjusting the main parameters of the model is to mitigate the risk of overfitting. Our team found that although some parameter settings achieved very high AUC scores locally, the scores on the leaderboard were not as impressive. This could be due to reduced model stability or overfitting, which resulted in no improvement in AUC on the test data. Since we couldn't pinpoint the exact cause, we believed that adjusting the parameters based on changes in the leaderboard scores was a safer and more reliable approach.</p>\n<h2>About the metric</h2>\n<p>As has been widely discussed recently, the metric trick can significantly improve leaderboard scores. Our team simulated a gradually decreasing weekly AUC and calculated the corresponding leaderboard scores. We found that lowering the AUC of approximately the first 20% of the data maximized the leaderboard scores.</p>\n<p>Before the 'This is the way' Notebook was released, our team used the following method to accurately reconstruct the dates.</p>\n<pre><code>def (regex_path):\n    chunks = []\n    for path in ((regex_path)):\n        exps = [\n            pl.().().(),\n            pl.().(pl.() == pl.().()).().(),#同一年份最大月份\n        ]\n        df = pl.(path).().(exps)\n        chunks.(df)\n\n    df = pl.(chunks, how=)\n    df = df.(subset=[])\n\n    df = df.()\n\n    df = df.(index = df.index[df[].()])\n    # df[].(, inplace=True)\n    df[].(, inplace=True)\n\n    df[] = df[].(int).(str)\n    df[] = df[].(int).(str)\n    df[] = pd.(df[] +  + df[], format=)\n    df.(columns=[,],inplace=True)\n    # df = df.(index=df.index[df[] == pd.()])\n\n    df.(,drop=True,inplace=True)\n\n    return df\nregex_path = TEST_DIR / \ndf_test_date = (regex_path)\ndf_test_date\n</code></pre>\n<p>However, 'This is the way' could more accurately infer the dates by calculating the similarity between the training and test data. As a result, our team ultimately adopted this method. Nevertheless, since the organizers used a time split method to separate the private leaderboard data from the public leaderboard data, the 'This is the way' approach did not improve our score at all. On the contrary, due to the inaccuracy of the date hack, the highest score we achieved was 0.591.</p>\n<p>Our team speculates that if we had used the 'This is the way' method to lower the AUC of the first 50% or 60% of the data, it might have better avoided the trap set by the organizers.</p>",
  "messages": [
    {
      "id": 2846053,
      "postDate": "2024-05-30T20:49:23.577Z",
      "content": "<h2>Overview</h2>\n<p>Our team utilized an ensemble of two models, LightGBM and CatBoost, for training and prediction. Without employing any hacks, our team achieved a best private leaderboard score of 0.524 (which was our final best score). With the use of a date hack, our team achieved a best private leaderboard score of 0.591.</p>\n<p>Notebook link: <a href=\"https://www.kaggle.com/code/arthurroland/home-credit-lgb-cat-date-hack-solution\" target=\"_blank\">https://www.kaggle.com/code/arthurroland/home-credit-lgb-cat-date-hack-solution</a></p>\n<h2>Feature engineering</h2>\n<p>For the LightGBM model, we selected 610 features, while for the CatBoost model, we chose 535 features. </p>\n<p>Here are some feature processing methods summarized by our team that can improve model performance:</p>\n<ol>\n<li><p>Our team found that features with a missing rate between 0.9 and 0.98 contained a significant amount of useful information. By adjusting the filtering threshold from 0.9 to 0.98, we were able to significantly improve the model's performance (from an initial score of 0.585 to 0.594). </p></li>\n<li><p>We processed the date features by creating a new feature representing the number of days between the given date and January 1, 1900, which also significantly enhanced the model's performance (improving the LightGBM single model score from 0.580 to 0.585).</p></li>\n<li><p>For string-type features with more than 200 unique categories, we retained the top 20 most frequent values and set the remaining values to null. Additionally, string-type features with more than 10,000 unique categories were primarily names, which we deemed to have no value. Therefore, we directly removed these features.</p></li>\n</ol>\n<p>Aside from the methods mentioned above, our team also performed minor feature processing. However, the improvements gained were minimal and can therefore be disregarded.</p>\n<h2>Hyperparameter tuning</h2>\n<p>Our team manually set the main parameters of the model, such as depth and num_leaves. The remaining parameters were adjusted using Optuna, and we ultimately obtained the following parameter settings.</p>\n<p>The reason for manually adjusting the main parameters of the model is to mitigate the risk of overfitting. Our team found that although some parameter settings achieved very high AUC scores locally, the scores on the leaderboard were not as impressive. This could be due to reduced model stability or overfitting, which resulted in no improvement in AUC on the test data. Since we couldn't pinpoint the exact cause, we believed that adjusting the parameters based on changes in the leaderboard scores was a safer and more reliable approach.</p>\n<h2>About the metric</h2>\n<p>As has been widely discussed recently, the metric trick can significantly improve leaderboard scores. Our team simulated a gradually decreasing weekly AUC and calculated the corresponding leaderboard scores. We found that lowering the AUC of approximately the first 20% of the data maximized the leaderboard scores.</p>\n<p>Before the 'This is the way' Notebook was released, our team used the following method to accurately reconstruct the dates.</p>\n<pre><code>def (regex_path):\n    chunks = []\n    for path in ((regex_path)):\n        exps = [\n            pl.().().(),\n            pl.().(pl.() == pl.().()).().(),#同一年份最大月份\n        ]\n        df = pl.(path).().(exps)\n        chunks.(df)\n\n    df = pl.(chunks, how=)\n    df = df.(subset=[])\n\n    df = df.()\n\n    df = df.(index = df.index[df[].()])\n    # df[].(, inplace=True)\n    df[].(, inplace=True)\n\n    df[] = df[].(int).(str)\n    df[] = df[].(int).(str)\n    df[] = pd.(df[] +  + df[], format=)\n    df.(columns=[,],inplace=True)\n    # df = df.(index=df.index[df[] == pd.()])\n\n    df.(,drop=True,inplace=True)\n\n    return df\nregex_path = TEST_DIR / \ndf_test_date = (regex_path)\ndf_test_date\n</code></pre>\n<p>However, 'This is the way' could more accurately infer the dates by calculating the similarity between the training and test data. As a result, our team ultimately adopted this method. Nevertheless, since the organizers used a time split method to separate the private leaderboard data from the public leaderboard data, the 'This is the way' approach did not improve our score at all. On the contrary, due to the inaccuracy of the date hack, the highest score we achieved was 0.591.</p>\n<p>Our team speculates that if we had used the 'This is the way' method to lower the AUC of the first 50% or 60% of the data, it might have better avoided the trap set by the organizers.</p>",
      "rawMarkdown": "## Overview\nOur team utilized an ensemble of two models, LightGBM and CatBoost, for training and prediction. Without employing any hacks, our team achieved a best private leaderboard score of 0.524 (which was our final best score). With the use of a date hack, our team achieved a best private leaderboard score of 0.591.\n\nNotebook link: https://www.kaggle.com/code/arthurroland/home-credit-lgb-cat-date-hack-solution\n\n## Feature engineering\nFor the LightGBM model, we selected 610 features, while for the CatBoost model, we chose 535 features. \n\nHere are some feature processing methods summarized by our team that can improve model performance:\n\n1. Our team found that features with a missing rate between 0.9 and 0.98 contained a significant amount of useful information. By adjusting the filtering threshold from 0.9 to 0.98, we were able to significantly improve the model's performance (from an initial score of 0.585 to 0.594). \n\n2. We processed the date features by creating a new feature representing the number of days between the given date and January 1, 1900, which also significantly enhanced the model's performance (improving the LightGBM single model score from 0.580 to 0.585).\n\n3. For string-type features with more than 200 unique categories, we retained the top 20 most frequent values and set the remaining values to null. Additionally, string-type features with more than 10,000 unique categories were primarily names, which we deemed to have no value. Therefore, we directly removed these features.\n\nAside from the methods mentioned above, our team also performed minor feature processing. However, the improvements gained were minimal and can therefore be disregarded.\n\n## Hyperparameter tuning\nOur team manually set the main parameters of the model, such as depth and num_leaves. The remaining parameters were adjusted using Optuna, and we ultimately obtained the following parameter settings.\n\nThe reason for manually adjusting the main parameters of the model is to mitigate the risk of overfitting. Our team found that although some parameter settings achieved very high AUC scores locally, the scores on the leaderboard were not as impressive. This could be due to reduced model stability or overfitting, which resulted in no improvement in AUC on the test data. Since we couldn't pinpoint the exact cause, we believed that adjusting the parameters based on changes in the leaderboard scores was a safer and more reliable approach.\n\n## About the metric\nAs has been widely discussed recently, the metric trick can significantly improve leaderboard scores. Our team simulated a gradually decreasing weekly AUC and calculated the corresponding leaderboard scores. We found that lowering the AUC of approximately the first 20% of the data maximized the leaderboard scores.\n\nBefore the 'This is the way' Notebook was released, our team used the following method to accurately reconstruct the dates.\n\n```\ndef getDate(regex_path):\n    chunks = []\n    for path in glob(str(regex_path)):\n        exps = [\n            pl.col(\"dpdmaxdateyear_596T\").max().alias(\"year\"),\n            pl.col(\"dpdmaxdatemonth_89T\").filter(pl.col(\"dpdmaxdateyear_596T\") == pl.col(\"dpdmaxdateyear_596T\").max()).max().alias(\"month\"),#同一年份最大月份\n        ]\n        df = pl.read_parquet(path).group_by(\"case_id\").agg(exps)\n        chunks.append(df)\n    \n    df = pl.concat(chunks, how=\"vertical_relaxed\")\n    df = df.unique(subset=[\"case_id\"])\n\n    df = df.to_pandas()\n    \n    df = df.drop(index = df.index[df[\"year\"].isna()])\n    # df[\"year\"].fillna(\"2019\", inplace=True)\n    df[\"month\"].fillna(\"12\", inplace=True)\n\n    df[\"year\"] = df[\"year\"].astype(int).astype(str)\n    df[\"month\"] = df[\"month\"].astype(int).astype(str)\n    df[\"datetime\"] = pd.to_datetime(df[\"year\"] + \"-\" + df[\"month\"], format=\"%Y-%m\")\n    df.drop(columns=[\"year\",\"month\"],inplace=True)\n    # df = df.drop(index=df.index[df[\"datetime\"] == pd.to_datetime(\"2019-01-01\")])\n    \n    df.set_index(\"case_id\",drop=True,inplace=True)\n    \n    return df\nregex_path = TEST_DIR / \"test_credit_bureau_a_1_*.parquet\"\ndf_test_date = getDate(regex_path)\ndf_test_date\n```\n\n\nHowever, 'This is the way' could more accurately infer the dates by calculating the similarity between the training and test data. As a result, our team ultimately adopted this method. Nevertheless, since the organizers used a time split method to separate the private leaderboard data from the public leaderboard data, the 'This is the way' approach did not improve our score at all. On the contrary, due to the inaccuracy of the date hack, the highest score we achieved was 0.591.\n\nOur team speculates that if we had used the 'This is the way' method to lower the AUC of the first 50% or 60% of the data, it might have better avoided the trap set by the organizers.",
      "votes": 15
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2846053": "## Overview\nOur team utilized an ensemble of two models, LightGBM and CatBoost, for training and prediction. Without employing any hacks, our team achieved a best private leaderboard score of 0.524 (which was our final best score). With the use of a date hack, our team achieved a best private leaderboard score of 0.591.\n\nNotebook link: https://www.kaggle.com/code/arthurroland/home-credit-lgb-cat-date-hack-solution\n\n## Feature engineering\nFor the LightGBM model, we selected 610 features, while for the CatBoost model, we chose 535 features. \n\nHere are some feature processing methods summarized by our team that can improve model performance:\n\n1. Our team found that features with a missing rate between 0.9 and 0.98 contained a significant amount of useful information. By adjusting the filtering threshold from 0.9 to 0.98, we were able to significantly improve the model's performance (from an initial score of 0.585 to 0.594). \n\n2. We processed the date features by creating a new feature representing the number of days between the given date and January 1, 1900, which also significantly enhanced the model's performance (improving the LightGBM single model score from 0.580 to 0.585).\n\n3. For string-type features with more than 200 unique categories, we retained the top 20 most frequent values and set the remaining values to null. Additionally, string-type features with more than 10,000 unique categories were primarily names, which we deemed to have no value. Therefore, we directly removed these features.\n\nAside from the methods mentioned above, our team also performed minor feature processing. However, the improvements gained were minimal and can therefore be disregarded.\n\n## Hyperparameter tuning\nOur team manually set the main parameters of the model, such as depth and num_leaves. The remaining parameters were adjusted using Optuna, and we ultimately obtained the following parameter settings.\n\nThe reason for manually adjusting the main parameters of the model is to mitigate the risk of overfitting. Our team found that although some parameter settings achieved very high AUC scores locally, the scores on the leaderboard were not as impressive. This could be due to reduced model stability or overfitting, which resulted in no improvement in AUC on the test data. Since we couldn't pinpoint the exact cause, we believed that adjusting the parameters based on changes in the leaderboard scores was a safer and more reliable approach.\n\n## About the metric\nAs has been widely discussed recently, the metric trick can significantly improve leaderboard scores. Our team simulated a gradually decreasing weekly AUC and calculated the corresponding leaderboard scores. We found that lowering the AUC of approximately the first 20% of the data maximized the leaderboard scores.\n\nBefore the 'This is the way' Notebook was released, our team used the following method to accurately reconstruct the dates.\n\n```\ndef getDate(regex_path):\n    chunks = []\n    for path in glob(str(regex_path)):\n        exps = [\n            pl.col(\"dpdmaxdateyear_596T\").max().alias(\"year\"),\n            pl.col(\"dpdmaxdatemonth_89T\").filter(pl.col(\"dpdmaxdateyear_596T\") == pl.col(\"dpdmaxdateyear_596T\").max()).max().alias(\"month\"),#同一年份最大月份\n        ]\n        df = pl.read_parquet(path).group_by(\"case_id\").agg(exps)\n        chunks.append(df)\n    \n    df = pl.concat(chunks, how=\"vertical_relaxed\")\n    df = df.unique(subset=[\"case_id\"])\n\n    df = df.to_pandas()\n    \n    df = df.drop(index = df.index[df[\"year\"].isna()])\n    # df[\"year\"].fillna(\"2019\", inplace=True)\n    df[\"month\"].fillna(\"12\", inplace=True)\n\n    df[\"year\"] = df[\"year\"].astype(int).astype(str)\n    df[\"month\"] = df[\"month\"].astype(int).astype(str)\n    df[\"datetime\"] = pd.to_datetime(df[\"year\"] + \"-\" + df[\"month\"], format=\"%Y-%m\")\n    df.drop(columns=[\"year\",\"month\"],inplace=True)\n    # df = df.drop(index=df.index[df[\"datetime\"] == pd.to_datetime(\"2019-01-01\")])\n    \n    df.set_index(\"case_id\",drop=True,inplace=True)\n    \n    return df\nregex_path = TEST_DIR / \"test_credit_bureau_a_1_*.parquet\"\ndf_test_date = getDate(regex_path)\ndf_test_date\n```\n\n\nHowever, 'This is the way' could more accurately infer the dates by calculating the similarity between the training and test data. As a result, our team ultimately adopted this method. Nevertheless, since the organizers used a time split method to separate the private leaderboard data from the public leaderboard data, the 'This is the way' approach did not improve our score at all. On the contrary, due to the inaccuracy of the date hack, the highest score we achieved was 0.591.\n\nOur team speculates that if we had used the 'This is the way' method to lower the AUC of the first 50% or 60% of the data, it might have better avoided the trap set by the organizers."
  }
}