{
  "id": 507959,
  "title": "Another simple metric hack can get  Public 0.626 / Private 0.616",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/507959",
  "author_name": "",
  "post_date": "2024-05-28T01:38:59.717850200Z",
  "votes": 23,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Thank you to the organizers and open-source contributors, you have brought me a lot of benefits。<br>\nUnfortunately, I did not select the high scoring code results for the B-list。<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20613639%2Fa055cda6b1b9234e5826d5e153decc5f%2F.png?generation=1716858659428039&amp;alt=media\" alt=\"another simple metric hack score\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20613639%2F0aaf100debd7b16bde01108316fae76b%2F2.png?generation=1716858835318399&amp;alt=media\" alt=\"another simple metric hack score\"></p>\n<p>Here, I would like to publicly disclose this method, which is probably the one chosen by the front row to maintain a stable ranking。Weeknum can be sorted through the dpdmaxdateyear_596T、dpdmaxdatemonth_89T、overdueamountmaxdateyear_2T、overdueamountmaxdatemonth_365T；<br>\nAdd the following post-processing on the basis of open-source 0.584 or 0.585 code，You can get a B-list score similar to the gold medal zone in the picture above；</p>\n<pre><code>def (regex_path, depth=None):\n    chunks = []\n    for path in ((regex_path)):\n        df = pl.(path)\n        df = df.(Pipeline.set_table_dtypes)\n        df = df.([, , ,\n                        ,,])\n        df = df.(df[] == )\n        df = df.([])\n        chunks.(df)\n    df = pl.(chunks, how=)\n    df = df.(subset=[])\n    return df\n\ndf_base =  (TEST_DIR / )\ntmp = (TEST_DIR / )\ntmp = df_base.(tmp, how=, on=)\ntmp = tmp.([, , ,\n                  ,]).()\n\ndf_test = df_test.(columns=[])\ndf_test = df_test.()\n\ny_pred = pd.(model.(df_test)[:, ], index=df_test.index)\n\ndf_subm = pd.(ROOT / )\ndf_subm = df_subm.()\ndf_subm[] = y_pred\n\ndf_subm = df_subm.(tmp,how=,on=[])\n\nyear_596T_ratio = df_subm[].().()/ (df_subm)\nyear_2T_ratio = df_subm[].().()/ (df_subm)\n\nif year_596T_ratio &lt;= year_2T_ratio:\n    df_subm = df_subm.(by=[, ])\nelse:\n    df_subm = df_subm.(by=[, ])\n\ndf_subm = df_subm.(drop=True)\nscore_col_index = df_subm.columns.()\ncut_off_index = ((df_subm) * )\ndf_subm.iloc[:cut_off_index, score_col_index] = (df_subm.iloc[:cut_off_index, score_col_index] - ).()\ndf_subm= df_subm.([,,\n                      ,],axis=)\ndf_subm = df_subm.()\ndf_subm.()\n</code></pre>\n<p>Although I didn't have a stable CV due to the influence of metrics, participating in this competition was more of a gain for me;<br>\nI will tidy up my mood and embark on the next kaggle journey. Thank you all, maybe we can meet again on Automated Essay Scoring 2.0.</p>",
  "messages": [
    {
      "id": "2840124",
      "postDate": "05/28/2024 01:38:59",
      "content": "<p>Thank you to the organizers and open-source contributors, you have brought me a lot of benefits。<br>\nUnfortunately, I did not select the high scoring code results for the B-list。<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20613639%2Fa055cda6b1b9234e5826d5e153decc5f%2F.png?generation=1716858659428039&amp;alt=media\" alt=\"another simple metric hack score\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20613639%2F0aaf100debd7b16bde01108316fae76b%2F2.png?generation=1716858835318399&amp;alt=media\" alt=\"another simple metric hack score\"></p>\n<p>Here, I would like to publicly disclose this method, which is probably the one chosen by the front row to maintain a stable ranking。Weeknum can be sorted through the dpdmaxdateyear_596T、dpdmaxdatemonth_89T、overdueamountmaxdateyear_2T、overdueamountmaxdatemonth_365T；<br>\nAdd the following post-processing on the basis of open-source 0.584 or 0.585 code，You can get a B-list score similar to the gold medal zone in the picture above；</p>\n<pre><code>def (regex_path, depth=None):\n    chunks = []\n    for path in ((regex_path)):\n        df = pl.(path)\n        df = df.(Pipeline.set_table_dtypes)\n        df = df.([, , ,\n                        ,,])\n        df = df.(df[] == )\n        df = df.([])\n        chunks.(df)\n    df = pl.(chunks, how=)\n    df = df.(subset=[])\n    return df\n\ndf_base =  (TEST_DIR / )\ntmp = (TEST_DIR / )\ntmp = df_base.(tmp, how=, on=)\ntmp = tmp.([, , ,\n                  ,]).()\n\ndf_test = df_test.(columns=[])\ndf_test = df_test.()\n\ny_pred = pd.(model.(df_test)[:, ], index=df_test.index)\n\ndf_subm = pd.(ROOT / )\ndf_subm = df_subm.()\ndf_subm[] = y_pred\n\ndf_subm = df_subm.(tmp,how=,on=[])\n\nyear_596T_ratio = df_subm[].().()/ (df_subm)\nyear_2T_ratio = df_subm[].().()/ (df_subm)\n\nif year_596T_ratio &lt;= year_2T_ratio:\n    df_subm = df_subm.(by=[, ])\nelse:\n    df_subm = df_subm.(by=[, ])\n\ndf_subm = df_subm.(drop=True)\nscore_col_index = df_subm.columns.()\ncut_off_index = ((df_subm) * )\ndf_subm.iloc[:cut_off_index, score_col_index] = (df_subm.iloc[:cut_off_index, score_col_index] - ).()\ndf_subm= df_subm.([,,\n                      ,],axis=)\ndf_subm = df_subm.()\ndf_subm.()\n</code></pre>\n<p>Although I didn't have a stable CV due to the influence of metrics, participating in this competition was more of a gain for me;<br>\nI will tidy up my mood and embark on the next kaggle journey. Thank you all, maybe we can meet again on Automated Essay Scoring 2.0.</p>",
      "rawMarkdown": "Thank you to the organizers and open-source contributors, you have brought me a lot of benefits。\nUnfortunately, I did not select the high scoring code results for the B-list。\n![another simple metric hack score](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20613639%2Fa055cda6b1b9234e5826d5e153decc5f%2F.png?generation=1716858659428039&alt=media)\n![another simple metric hack score](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20613639%2F0aaf100debd7b16bde01108316fae76b%2F2.png?generation=1716858835318399&alt=media)\n\nHere, I would like to publicly disclose this method, which is probably the one chosen by the front row to maintain a stable ranking。Weeknum can be sorted through the dpdmaxdateyear_596T、dpdmaxdatemonth_89T、overdueamountmaxdateyear_2T、overdueamountmaxdatemonth_365T；\nAdd the following post-processing on the basis of open-source 0.584 or 0.585 code，You can get a B-list score similar to the gold medal zone in the picture above；\n```\ndef read_files2(regex_path, depth=None):\n    chunks = []\n    for path in glob(str(regex_path)):\n        df = pl.read_parquet(path)\n        df = df.pipe(Pipeline.set_table_dtypes)\n        df = df.select(['case_id', 'dpdmaxdateyear_596T', 'dpdmaxdatemonth_89T',\n                        'overdueamountmaxdateyear_2T','overdueamountmaxdatemonth_365T','num_group1'])\n        df = df.filter(df['num_group1'] == 0)\n        df = df.drop(['num_group1'])\n        chunks.append(df)\n    df = pl.concat(chunks, how=\"vertical_relaxed\")\n    df = df.unique(subset=[\"case_id\"])\n    return df\n\ndf_base =  read_file(TEST_DIR / \"test_base.parquet\")\ntmp = read_files2(TEST_DIR / \"test_credit_bureau_a_1_*.parquet\")\ntmp = df_base.join(tmp, how=\"left\", on=\"case_id\")\ntmp = tmp.select(['case_id', 'dpdmaxdateyear_596T', 'dpdmaxdatemonth_89T',\n                  'overdueamountmaxdateyear_2T','overdueamountmaxdatemonth_365T']).to_pandas()\n\ndf_test = df_test.drop(columns=[\"WEEK_NUM\"])\ndf_test = df_test.set_index(\"case_id\")\n\ny_pred = pd.Series(model.predict_proba(df_test)[:, 1], index=df_test.index)\n\ndf_subm = pd.read_csv(ROOT / \"sample_submission.csv\")\ndf_subm = df_subm.set_index(\"case_id\")\ndf_subm[\"score\"] = y_pred\n\ndf_subm = df_subm.merge(tmp,how='left',on=['case_id'])\n\nyear_596T_ratio = df_subm['dpdmaxdateyear_596T'].isnull().sum()/ len(df_subm)\nyear_2T_ratio = df_subm['overdueamountmaxdateyear_2T'].isnull().sum()/ len(df_subm)\n\nif year_596T_ratio <= year_2T_ratio:\n    df_subm = df_subm.sort_values(by=['dpdmaxdateyear_596T', 'dpdmaxdatemonth_89T'])\nelse:\n    df_subm = df_subm.sort_values(by=['overdueamountmaxdateyear_2T', 'overdueamountmaxdatemonth_365T'])\n    \ndf_subm = df_subm.reset_index(drop=True)\nscore_col_index = df_subm.columns.get_loc('score')\ncut_off_index = int(len(df_subm) * 0.35)\ndf_subm.iloc[:cut_off_index, score_col_index] = (df_subm.iloc[:cut_off_index, score_col_index] - 0.055).clip(0)\ndf_subm= df_subm.drop(['dpdmaxdateyear_596T','dpdmaxdatemonth_89T',\n                      'overdueamountmaxdateyear_2T','overdueamountmaxdatemonth_365T'],axis=1)\ndf_subm = df_subm.set_index(\"case_id\")\ndf_subm.to_csv(\"submission.csv\")\n```\n\nAlthough I didn't have a stable CV due to the influence of metrics, participating in this competition was more of a gain for me;\nI will tidy up my mood and embark on the next kaggle journey. Thank you all, maybe we can meet again on Automated Essay Scoring 2.0.",
      "votes": null
    },
    {
      "id": "2840152",
      "postDate": "05/28/2024 01:53:00",
      "content": "<p>Wonderful!</p>",
      "rawMarkdown": "Wonderful!",
      "votes": null
    },
    {
      "id": "2840179",
      "postDate": "05/28/2024 02:22:52",
      "content": "<p>太可惜了，大佬！What a pity</p>",
      "rawMarkdown": "太可惜了，大佬！What a pity",
      "votes": null
    },
    {
      "id": "2840194",
      "postDate": "05/28/2024 02:38:22",
      "content": "<p>Awesome !!!</p>",
      "rawMarkdown": "Awesome !!!",
      "votes": null
    },
    {
      "id": "2840646",
      "postDate": "05/28/2024 07:26:37",
      "content": "<p>Love youuuuu!!!!!!!!!!! &lt;3. </p>",
      "rawMarkdown": "Love youuuuu!!!!!!!!!!! <3.",
      "votes": null
    },
    {
      "id": "2841655",
      "postDate": "05/28/2024 16:36:30",
      "content": "<p>确实牛 实在太可惜了</p>",
      "rawMarkdown": "确实牛 实在太可惜了",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2840152,
      "author_name": "polsyuus",
      "author_url": "",
      "post_date": "05/28/2024 01:53:00",
      "content": "<p>Wonderful!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2840179,
      "author_name": "wessonz",
      "author_url": "",
      "post_date": "05/28/2024 02:22:52",
      "content": "<p>太可惜了，大佬！What a pity</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2840194,
      "author_name": "nguyenosaurus",
      "author_url": "",
      "post_date": "05/28/2024 02:38:22",
      "content": "<p>Awesome !!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2840646,
      "author_name": "marcyyl",
      "author_url": "",
      "post_date": "05/28/2024 07:26:37",
      "content": "<p>Love youuuuu!!!!!!!!!!! &lt;3. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2841655,
      "author_name": "licongdu",
      "author_url": "",
      "post_date": "05/28/2024 16:36:30",
      "content": "<p>确实牛 实在太可惜了</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2840124": "Thank you to the organizers and open-source contributors, you have brought me a lot of benefits。\nUnfortunately, I did not select the high scoring code results for the B-list。\n![another simple metric hack score](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20613639%2Fa055cda6b1b9234e5826d5e153decc5f%2F.png?generation=1716858659428039&alt=media)\n![another simple metric hack score](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20613639%2F0aaf100debd7b16bde01108316fae76b%2F2.png?generation=1716858835318399&alt=media)\n\nHere, I would like to publicly disclose this method, which is probably the one chosen by the front row to maintain a stable ranking。Weeknum can be sorted through the dpdmaxdateyear_596T、dpdmaxdatemonth_89T、overdueamountmaxdateyear_2T、overdueamountmaxdatemonth_365T；\nAdd the following post-processing on the basis of open-source 0.584 or 0.585 code，You can get a B-list score similar to the gold medal zone in the picture above；\n```\ndef read_files2(regex_path, depth=None):\n    chunks = []\n    for path in glob(str(regex_path)):\n        df = pl.read_parquet(path)\n        df = df.pipe(Pipeline.set_table_dtypes)\n        df = df.select(['case_id', 'dpdmaxdateyear_596T', 'dpdmaxdatemonth_89T',\n                        'overdueamountmaxdateyear_2T','overdueamountmaxdatemonth_365T','num_group1'])\n        df = df.filter(df['num_group1'] == 0)\n        df = df.drop(['num_group1'])\n        chunks.append(df)\n    df = pl.concat(chunks, how=\"vertical_relaxed\")\n    df = df.unique(subset=[\"case_id\"])\n    return df\n\ndf_base =  read_file(TEST_DIR / \"test_base.parquet\")\ntmp = read_files2(TEST_DIR / \"test_credit_bureau_a_1_*.parquet\")\ntmp = df_base.join(tmp, how=\"left\", on=\"case_id\")\ntmp = tmp.select(['case_id', 'dpdmaxdateyear_596T', 'dpdmaxdatemonth_89T',\n                  'overdueamountmaxdateyear_2T','overdueamountmaxdatemonth_365T']).to_pandas()\n\ndf_test = df_test.drop(columns=[\"WEEK_NUM\"])\ndf_test = df_test.set_index(\"case_id\")\n\ny_pred = pd.Series(model.predict_proba(df_test)[:, 1], index=df_test.index)\n\ndf_subm = pd.read_csv(ROOT / \"sample_submission.csv\")\ndf_subm = df_subm.set_index(\"case_id\")\ndf_subm[\"score\"] = y_pred\n\ndf_subm = df_subm.merge(tmp,how='left',on=['case_id'])\n\nyear_596T_ratio = df_subm['dpdmaxdateyear_596T'].isnull().sum()/ len(df_subm)\nyear_2T_ratio = df_subm['overdueamountmaxdateyear_2T'].isnull().sum()/ len(df_subm)\n\nif year_596T_ratio <= year_2T_ratio:\n    df_subm = df_subm.sort_values(by=['dpdmaxdateyear_596T', 'dpdmaxdatemonth_89T'])\nelse:\n    df_subm = df_subm.sort_values(by=['overdueamountmaxdateyear_2T', 'overdueamountmaxdatemonth_365T'])\n    \ndf_subm = df_subm.reset_index(drop=True)\nscore_col_index = df_subm.columns.get_loc('score')\ncut_off_index = int(len(df_subm) * 0.35)\ndf_subm.iloc[:cut_off_index, score_col_index] = (df_subm.iloc[:cut_off_index, score_col_index] - 0.055).clip(0)\ndf_subm= df_subm.drop(['dpdmaxdateyear_596T','dpdmaxdatemonth_89T',\n                      'overdueamountmaxdateyear_2T','overdueamountmaxdatemonth_365T'],axis=1)\ndf_subm = df_subm.set_index(\"case_id\")\ndf_subm.to_csv(\"submission.csv\")\n```\n\nAlthough I didn't have a stable CV due to the influence of metrics, participating in this competition was more of a gain for me;\nI will tidy up my mood and embark on the next kaggle journey. Thank you all, maybe we can meet again on Automated Essay Scoring 2.0.",
    "2840152": "Wonderful!",
    "2840179": "太可惜了，大佬！What a pity",
    "2840194": "Awesome !!!",
    "2840646": "Love youuuuu!!!!!!!!!!! <3.",
    "2841655": "确实牛 实在太可惜了"
  },
  "source": "meta"
}