{
  "id": 508343,
  "title": "70th place solution - no metric hack",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/508343",
  "author_name": "LucaMTB",
  "post_date": "2024-05-29T08:57:31.567000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>Overview</h1>\n<p>We used data from all table and in total a number of 442 feature. For bureau_a_1 files we justed use columns for closed contracts. </p>\n<p><code>col_closed_cont = ['case_id',\n'annualeffectiverate_199L',\n'classificationofcontr_400M',\n'contractst_964M',\n'contractsum_5085717L',\n'credlmt_230A',\n'dateofcredend_353D',\n'dateofcredstart_739D',\n'dateofrealrepmt_138D',\n'debtoverdue_47A',\n'description_351M',\n'dpdmax_757P',\n'financialinstitution_382M',\n'instlamount_852A',\n'interestrate_508L',\n'lastupdate_388D',\n'monthlyinstlamount_674A',\n'nominalrate_498L',\n'num_group1',\n'numberofcontrsvalue_358L',\n'numberofinstls_229L',\n'numberofoutstandinstls_520L',\n'numberofoverdueinstlmax_1151L',\n'numberofoverdueinstlmaxdat_148D',\n'numberofoverdueinstls_834L',\n'outstandingamount_354A',\n'overdueamount_31A',\n'overdueamountmax2_398A',\n'overdueamountmax2date_1002D',\n'overdueamountmax_35A',\n'overdueamountmaxdatemonth_284T',\n'periodicityofpmts_1102L',\n'prolongationcount_1120L',\n'purposeofcred_874M',\n'refreshdate_3813885D',\n'residualamount_488A',\n'subjectrole_93M',\n'totalamount_6A',\n'totaldebtoverduevalue_718A',\n'totaloutstanddebtvalue_668A']</code></p>\n<p>Like most teams we used mean, max, min for aggregation.</p>\n<p>We used an ensemble of LightGBM, XGBoost, CatBoost Classifier. </p>\n<h1>Details</h1>\n<h2>Remove Data with drift</h2>\n<p>We made an extensive search for data with drift</p>\n<p><code>df_base = df_base.drop(\"mean_tenor_203L\" , \"mean_pmtnum_8L\",  \n                           \"eir_270L\",\"numberofqueries_373L\",\n       'firstquarter_103L', 'mindbddpdlast24m_3658935P'\n                          \"processingdate_168D\",\n                           \"postype_4733339M\",\n                           \"assignmentdate_238D\", \n                           \"pmts_year_1139T\" ,\n                           \"pmts_year_507T\" , \n                           \"subjectroles_name_541M\",\n                           \"collaterals_typeofguarante_359M\" ,\n                           \"firstclxcampaign_1125D\",\n                            \"revolvingaccount_394A\",\n                           'bankacctype_710L'                       \n                          )</code></p>\n<h2>Make model more robust through transformation</h2>\n<p>We transformed some columns to make the model more robust</p>\n<p><code>df_base = df_base.with_columns(       \n        (pl.col(\"mean_installmentamount_644A\")-\n         df_base.select(pl.min(\"mean_installmentamount_644A\"))).pow(2),\n        ((pl.col(\"pctinstlsallpaidearl3d_427L\")-\n         df_base.select(pl.min(\"pctinstlsallpaidearl3d_427L\")))*100.).pow(1.0),        \n         ((pl.col(\"pctinstlsallpaidlat10d_839L\")-\n         df_base.select(pl.min(\"pctinstlsallpaidlat10d_839L\")))*100.).pow(1.0),      \n         ((pl.col(\"pctinstlsallpaidlate1d_3546856L\")-\n         df_base.select(pl.min(\"pctinstlsallpaidlate1d_3546856L\")))*100.).pow(1.0),        \n         ((pl.col(\"pctinstlsallpaidlate4d_3546849L\")-\n         df_base.select(pl.min(\"pctinstlsallpaidlate4d_3546849L\")))*100.).pow(1.0),        \n        ((pl.col(\"pctinstlsallpaidlate6d_3546844L\")-\n         df_base.select(pl.min(\"pctinstlsallpaidlate6d_3546844L\")))*100.).pow(1.0),         \n        (pl.col(\"mean_amount_1115A\")-\n         df_base.select(pl.min(\"mean_amount_1115A\"))).pow(2),               \n    )</code></p>\n<h3>RobustScaler and QuantileTransfomer</h3>\n<p><code>transformer = RobustScaler()\nqt = QuantileTransformer(random_state=0,n_quantiles=500)</code></p>\n<p><code>list_columns_to_transform = ['mean_processingdate_168D']</code></p>\n<p><code>list_columns_to_qt_transform= ['mean_revolvingaccount_394A','maxinstallast24m_3658928A','maxpmtlast3m_4525190A',                'lastapprcredamount_781A']</code></p>\n<p><code>for col in df_train.columns:\n    if \"P\" in col[-1]:\n        list_columns_to_transform.append(col)</code></p>\n<p><code>transformer.fit_transform(df_train[list_columns_to_transform])\nqt.fit_transform(df_train[list_columns_to_qt_transform])</code></p>\n<h1>Cross-Validation</h1>\n<p>We trained on the first 10 weeks and validated metrics on remaining 80 weeks. There was a slitly correlation to the leaderboard. It helped in some case. Both like most teams we didnt found a realy good correlation</p>",
  "messages": [
    {
      "id": 2842778,
      "postDate": "2024-05-29T08:57:31.567Z",
      "content": "<h1>Overview</h1>\n<p>We used data from all table and in total a number of 442 feature. For bureau_a_1 files we justed use columns for closed contracts. </p>\n<p><code>col_closed_cont = ['case_id',\n'annualeffectiverate_199L',\n'classificationofcontr_400M',\n'contractst_964M',\n'contractsum_5085717L',\n'credlmt_230A',\n'dateofcredend_353D',\n'dateofcredstart_739D',\n'dateofrealrepmt_138D',\n'debtoverdue_47A',\n'description_351M',\n'dpdmax_757P',\n'financialinstitution_382M',\n'instlamount_852A',\n'interestrate_508L',\n'lastupdate_388D',\n'monthlyinstlamount_674A',\n'nominalrate_498L',\n'num_group1',\n'numberofcontrsvalue_358L',\n'numberofinstls_229L',\n'numberofoutstandinstls_520L',\n'numberofoverdueinstlmax_1151L',\n'numberofoverdueinstlmaxdat_148D',\n'numberofoverdueinstls_834L',\n'outstandingamount_354A',\n'overdueamount_31A',\n'overdueamountmax2_398A',\n'overdueamountmax2date_1002D',\n'overdueamountmax_35A',\n'overdueamountmaxdatemonth_284T',\n'periodicityofpmts_1102L',\n'prolongationcount_1120L',\n'purposeofcred_874M',\n'refreshdate_3813885D',\n'residualamount_488A',\n'subjectrole_93M',\n'totalamount_6A',\n'totaldebtoverduevalue_718A',\n'totaloutstanddebtvalue_668A']</code></p>\n<p>Like most teams we used mean, max, min for aggregation.</p>\n<p>We used an ensemble of LightGBM, XGBoost, CatBoost Classifier. </p>\n<h1>Details</h1>\n<h2>Remove Data with drift</h2>\n<p>We made an extensive search for data with drift</p>\n<p><code>df_base = df_base.drop(\"mean_tenor_203L\" , \"mean_pmtnum_8L\",  \n                           \"eir_270L\",\"numberofqueries_373L\",\n       'firstquarter_103L', 'mindbddpdlast24m_3658935P'\n                          \"processingdate_168D\",\n                           \"postype_4733339M\",\n                           \"assignmentdate_238D\", \n                           \"pmts_year_1139T\" ,\n                           \"pmts_year_507T\" , \n                           \"subjectroles_name_541M\",\n                           \"collaterals_typeofguarante_359M\" ,\n                           \"firstclxcampaign_1125D\",\n                            \"revolvingaccount_394A\",\n                           'bankacctype_710L'                       \n                          )</code></p>\n<h2>Make model more robust through transformation</h2>\n<p>We transformed some columns to make the model more robust</p>\n<p><code>df_base = df_base.with_columns(       \n        (pl.col(\"mean_installmentamount_644A\")-\n         df_base.select(pl.min(\"mean_installmentamount_644A\"))).pow(2),\n        ((pl.col(\"pctinstlsallpaidearl3d_427L\")-\n         df_base.select(pl.min(\"pctinstlsallpaidearl3d_427L\")))*100.).pow(1.0),        \n         ((pl.col(\"pctinstlsallpaidlat10d_839L\")-\n         df_base.select(pl.min(\"pctinstlsallpaidlat10d_839L\")))*100.).pow(1.0),      \n         ((pl.col(\"pctinstlsallpaidlate1d_3546856L\")-\n         df_base.select(pl.min(\"pctinstlsallpaidlate1d_3546856L\")))*100.).pow(1.0),        \n         ((pl.col(\"pctinstlsallpaidlate4d_3546849L\")-\n         df_base.select(pl.min(\"pctinstlsallpaidlate4d_3546849L\")))*100.).pow(1.0),        \n        ((pl.col(\"pctinstlsallpaidlate6d_3546844L\")-\n         df_base.select(pl.min(\"pctinstlsallpaidlate6d_3546844L\")))*100.).pow(1.0),         \n        (pl.col(\"mean_amount_1115A\")-\n         df_base.select(pl.min(\"mean_amount_1115A\"))).pow(2),               \n    )</code></p>\n<h3>RobustScaler and QuantileTransfomer</h3>\n<p><code>transformer = RobustScaler()\nqt = QuantileTransformer(random_state=0,n_quantiles=500)</code></p>\n<p><code>list_columns_to_transform = ['mean_processingdate_168D']</code></p>\n<p><code>list_columns_to_qt_transform= ['mean_revolvingaccount_394A','maxinstallast24m_3658928A','maxpmtlast3m_4525190A',                'lastapprcredamount_781A']</code></p>\n<p><code>for col in df_train.columns:\n    if \"P\" in col[-1]:\n        list_columns_to_transform.append(col)</code></p>\n<p><code>transformer.fit_transform(df_train[list_columns_to_transform])\nqt.fit_transform(df_train[list_columns_to_qt_transform])</code></p>\n<h1>Cross-Validation</h1>\n<p>We trained on the first 10 weeks and validated metrics on remaining 80 weeks. There was a slitly correlation to the leaderboard. It helped in some case. Both like most teams we didnt found a realy good correlation</p>",
      "rawMarkdown": "# Overview\n\nWe used data from all table and in total a number of 442 feature. For bureau_a_1 files we justed use columns for closed contracts. \n\n`col_closed_cont = ['case_id',\n'annualeffectiverate_199L',\n'classificationofcontr_400M',\n'contractst_964M',\n'contractsum_5085717L',\n'credlmt_230A',\n'dateofcredend_353D',\n'dateofcredstart_739D',\n'dateofrealrepmt_138D',\n'debtoverdue_47A',\n'description_351M',\n'dpdmax_757P',\n'financialinstitution_382M',\n'instlamount_852A',\n'interestrate_508L',\n'lastupdate_388D',\n'monthlyinstlamount_674A',\n'nominalrate_498L',\n'num_group1',\n'numberofcontrsvalue_358L',\n'numberofinstls_229L',\n'numberofoutstandinstls_520L',\n'numberofoverdueinstlmax_1151L',\n'numberofoverdueinstlmaxdat_148D',\n'numberofoverdueinstls_834L',\n'outstandingamount_354A',\n'overdueamount_31A',\n'overdueamountmax2_398A',\n'overdueamountmax2date_1002D',\n'overdueamountmax_35A',\n'overdueamountmaxdatemonth_284T',\n'periodicityofpmts_1102L',\n'prolongationcount_1120L',\n'purposeofcred_874M',\n'refreshdate_3813885D',\n'residualamount_488A',\n'subjectrole_93M',\n'totalamount_6A',\n'totaldebtoverduevalue_718A',\n'totaloutstanddebtvalue_668A']`\n\nLike most teams we used mean, max, min for aggregation.\n\n\nWe used an ensemble of LightGBM, XGBoost, CatBoost Classifier. \n\n# Details\n\n## Remove Data with drift\n\nWe made an extensive search for data with drift\n\n`df_base = df_base.drop(\"mean_tenor_203L\" , \"mean_pmtnum_8L\",  \n                           \"eir_270L\",\"numberofqueries_373L\",\n       'firstquarter_103L', 'mindbddpdlast24m_3658935P'\n                          \"processingdate_168D\",\n                           \"postype_4733339M\",\n                           \"assignmentdate_238D\", \n                           \"pmts_year_1139T\" ,\n                           \"pmts_year_507T\" , \n                           \"subjectroles_name_541M\",\n                           \"collaterals_typeofguarante_359M\" ,\n                           \"firstclxcampaign_1125D\",\n                            \"revolvingaccount_394A\",\n                           'bankacctype_710L'                       \n                          )`\n\n\n## Make model more robust through transformation\n\nWe transformed some columns to make the model more robust\n\n`df_base = df_base.with_columns(       \n        (pl.col(\"mean_installmentamount_644A\")-\n         df_base.select(pl.min(\"mean_installmentamount_644A\"))).pow(2),\n        ((pl.col(\"pctinstlsallpaidearl3d_427L\")-\n         df_base.select(pl.min(\"pctinstlsallpaidearl3d_427L\")))*100.).pow(1.0),        \n         ((pl.col(\"pctinstlsallpaidlat10d_839L\")-\n         df_base.select(pl.min(\"pctinstlsallpaidlat10d_839L\")))*100.).pow(1.0),      \n         ((pl.col(\"pctinstlsallpaidlate1d_3546856L\")-\n         df_base.select(pl.min(\"pctinstlsallpaidlate1d_3546856L\")))*100.).pow(1.0),        \n         ((pl.col(\"pctinstlsallpaidlate4d_3546849L\")-\n         df_base.select(pl.min(\"pctinstlsallpaidlate4d_3546849L\")))*100.).pow(1.0),        \n        ((pl.col(\"pctinstlsallpaidlate6d_3546844L\")-\n         df_base.select(pl.min(\"pctinstlsallpaidlate6d_3546844L\")))*100.).pow(1.0),         \n        (pl.col(\"mean_amount_1115A\")-\n         df_base.select(pl.min(\"mean_amount_1115A\"))).pow(2),               \n    )`\n\n### RobustScaler and QuantileTransfomer\n\n`transformer = RobustScaler()\nqt = QuantileTransformer(random_state=0,n_quantiles=500)`\n\n`list_columns_to_transform = ['mean_processingdate_168D'] `\n\n`list_columns_to_qt_transform= ['mean_revolvingaccount_394A','maxinstallast24m_3658928A','maxpmtlast3m_4525190A',                'lastapprcredamount_781A'] `\n\n`for col in df_train.columns:\n    if \"P\" in col[-1]:\n        list_columns_to_transform.append(col) `\n\n`transformer.fit_transform(df_train[list_columns_to_transform])\nqt.fit_transform(df_train[list_columns_to_qt_transform])`\n\n# Cross-Validation\n\nWe trained on the first 10 weeks and validated metrics on remaining 80 weeks. There was a slitly correlation to the leaderboard. It helped in some case. Both like most teams we didnt found a realy good correlation",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2842778": "# Overview\n\nWe used data from all table and in total a number of 442 feature. For bureau_a_1 files we justed use columns for closed contracts. \n\n`col_closed_cont = ['case_id',\n'annualeffectiverate_199L',\n'classificationofcontr_400M',\n'contractst_964M',\n'contractsum_5085717L',\n'credlmt_230A',\n'dateofcredend_353D',\n'dateofcredstart_739D',\n'dateofrealrepmt_138D',\n'debtoverdue_47A',\n'description_351M',\n'dpdmax_757P',\n'financialinstitution_382M',\n'instlamount_852A',\n'interestrate_508L',\n'lastupdate_388D',\n'monthlyinstlamount_674A',\n'nominalrate_498L',\n'num_group1',\n'numberofcontrsvalue_358L',\n'numberofinstls_229L',\n'numberofoutstandinstls_520L',\n'numberofoverdueinstlmax_1151L',\n'numberofoverdueinstlmaxdat_148D',\n'numberofoverdueinstls_834L',\n'outstandingamount_354A',\n'overdueamount_31A',\n'overdueamountmax2_398A',\n'overdueamountmax2date_1002D',\n'overdueamountmax_35A',\n'overdueamountmaxdatemonth_284T',\n'periodicityofpmts_1102L',\n'prolongationcount_1120L',\n'purposeofcred_874M',\n'refreshdate_3813885D',\n'residualamount_488A',\n'subjectrole_93M',\n'totalamount_6A',\n'totaldebtoverduevalue_718A',\n'totaloutstanddebtvalue_668A']`\n\nLike most teams we used mean, max, min for aggregation.\n\n\nWe used an ensemble of LightGBM, XGBoost, CatBoost Classifier. \n\n# Details\n\n## Remove Data with drift\n\nWe made an extensive search for data with drift\n\n`df_base = df_base.drop(\"mean_tenor_203L\" , \"mean_pmtnum_8L\",  \n                           \"eir_270L\",\"numberofqueries_373L\",\n       'firstquarter_103L', 'mindbddpdlast24m_3658935P'\n                          \"processingdate_168D\",\n                           \"postype_4733339M\",\n                           \"assignmentdate_238D\", \n                           \"pmts_year_1139T\" ,\n                           \"pmts_year_507T\" , \n                           \"subjectroles_name_541M\",\n                           \"collaterals_typeofguarante_359M\" ,\n                           \"firstclxcampaign_1125D\",\n                            \"revolvingaccount_394A\",\n                           'bankacctype_710L'                       \n                          )`\n\n\n## Make model more robust through transformation\n\nWe transformed some columns to make the model more robust\n\n`df_base = df_base.with_columns(       \n        (pl.col(\"mean_installmentamount_644A\")-\n         df_base.select(pl.min(\"mean_installmentamount_644A\"))).pow(2),\n        ((pl.col(\"pctinstlsallpaidearl3d_427L\")-\n         df_base.select(pl.min(\"pctinstlsallpaidearl3d_427L\")))*100.).pow(1.0),        \n         ((pl.col(\"pctinstlsallpaidlat10d_839L\")-\n         df_base.select(pl.min(\"pctinstlsallpaidlat10d_839L\")))*100.).pow(1.0),      \n         ((pl.col(\"pctinstlsallpaidlate1d_3546856L\")-\n         df_base.select(pl.min(\"pctinstlsallpaidlate1d_3546856L\")))*100.).pow(1.0),        \n         ((pl.col(\"pctinstlsallpaidlate4d_3546849L\")-\n         df_base.select(pl.min(\"pctinstlsallpaidlate4d_3546849L\")))*100.).pow(1.0),        \n        ((pl.col(\"pctinstlsallpaidlate6d_3546844L\")-\n         df_base.select(pl.min(\"pctinstlsallpaidlate6d_3546844L\")))*100.).pow(1.0),         \n        (pl.col(\"mean_amount_1115A\")-\n         df_base.select(pl.min(\"mean_amount_1115A\"))).pow(2),               \n    )`\n\n### RobustScaler and QuantileTransfomer\n\n`transformer = RobustScaler()\nqt = QuantileTransformer(random_state=0,n_quantiles=500)`\n\n`list_columns_to_transform = ['mean_processingdate_168D'] `\n\n`list_columns_to_qt_transform= ['mean_revolvingaccount_394A','maxinstallast24m_3658928A','maxpmtlast3m_4525190A',                'lastapprcredamount_781A'] `\n\n`for col in df_train.columns:\n    if \"P\" in col[-1]:\n        list_columns_to_transform.append(col) `\n\n`transformer.fit_transform(df_train[list_columns_to_transform])\nqt.fit_transform(df_train[list_columns_to_qt_transform])`\n\n# Cross-Validation\n\nWe trained on the first 10 weeks and validated metrics on remaining 80 weeks. There was a slitly correlation to the leaderboard. It helped in some case. Both like most teams we didnt found a realy good correlation"
  }
}