{
  "id": 507949,
  "title": "Custom Loss: Public 0.637 / Private 0.544",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/507949",
  "author_name": "gromml",
  "post_date": "2024-05-28T00:47:31.861000",
  "votes": 7,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Here is our work with the custom loss (called StableRocAucLoss):</p>\n<p><a href=\"https://www.kaggle.com/code/gromml/custom-loss-stablerocaucloss\" target=\"_blank\">https://www.kaggle.com/code/gromml/custom-loss-stablerocaucloss</a></p>\n<p>Will be interesting to see approaches of other participants to address stability!</p>",
  "messages": [
    {
      "id": 2840076,
      "postDate": "2024-05-28T00:47:31.863Z",
      "content": "<p>Here is our work with the custom loss (called StableRocAucLoss):</p>\n<p><a href=\"https://www.kaggle.com/code/gromml/custom-loss-stablerocaucloss\" target=\"_blank\">https://www.kaggle.com/code/gromml/custom-loss-stablerocaucloss</a></p>\n<p>Will be interesting to see approaches of other participants to address stability!</p>",
      "rawMarkdown": "Here is our work with the custom loss (called StableRocAucLoss):\n\nhttps://www.kaggle.com/code/gromml/custom-loss-stablerocaucloss\n\nWill be interesting to see approaches of other participants to address stability!",
      "votes": 7
    },
    {
      "id": 2840159,
      "postDate": "2024-05-28T02:02:08.193Z",
      "content": "<p>Good and interesting approach. But, you should specify that the .544 is with metric hacking.</p>",
      "rawMarkdown": "Good and interesting approach. But, you should specify that the .544 is with metric hacking.",
      "votes": 4,
      "replies": [
        {
          "id": 2840551,
          "postDate": "2024-05-28T06:20:04.750Z",
          "content": "<p>Yes, interesting approach. I've tried it without metric hack, and it showed not excellent but pretty good results: Public - 0.585, Private - 0.502.</p>",
          "rawMarkdown": "Yes, interesting approach. I've tried it without metric hack, and it showed not excellent but pretty good results: Public - 0.585, Private - 0.502.",
          "votes": 1,
          "replies": [
            {
              "id": 2840744,
              "postDate": "2024-05-28T08:43:44.273Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        },
        {
          "id": 2840738,
          "postDate": "2024-05-28T08:40:12.573Z",
          "content": "<p>Yes, 0.544 uses the metric trick. One can check the second version of the notebook, V2 does not utilize any metric tricks. It has 0.585 / 0.502 score.</p>",
          "rawMarkdown": "Yes, 0.544 uses the metric trick. One can check the second version of the notebook, V2 does not utilize any metric tricks. It has 0.585 / 0.502 score.\n\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 2869840,
      "postDate": "2024-06-13T09:57:04.173Z",
      "content": "<p>Also, there is a file below with some tests of our loss.</p>",
      "rawMarkdown": "Also, there is a file below with some tests of our loss.",
      "votes": 1
    },
    {
      "id": 2866618,
      "postDate": "2024-06-11T12:07:38.893Z",
      "content": "<p>The competition metric is given by the following equation:<br>\nstability metric = mean(gini) + 88 * min(0, a) - 0.5 * std(residuals).</p>\n<p>At the beginning of the competition, we wanted to construct a loss consisting of three terms: the first term should have addressed mean(gini), the second one could have been responsible for the time decay, and the last one should have addressed std of residuals. However, when <code>WEEK_NUM</code> was replaced with a constant, we decided that we could hardly incorporate decay (that is, the 2nd term) into our loss function without information about time and dates. That is why, we decided to focus on only two terms with the hope of minimizing the decay effect throughout the last term. Indeed, if there is no trend in gini scores, there is a high chance that std of gini scores is lower in comparison with a linear trend in gini scores. This is how we came to our custom loss described in our public notebook: <a href=\"https://www.kaggle.com/code/gromml/custom-loss-stablerocaucloss\" target=\"_blank\">https://www.kaggle.com/code/gromml/custom-loss-stablerocaucloss</a>. In this notebook one can find additional details about our solution (preprocessing, ensembling, etc.).</p>\n<p>As mentioned in this public notebook, our ensemble consists of 21 models:</p>\n<ul>\n<li>7 lightgbm models (one model for each of the 7 folds) trained using a default loss;</li>\n<li>7 catboost models trained using a default loss (one model for each fold);</li>\n<li>7 lightgbm models trained using the custom loss (which was introduced in the public notebook).</li>\n</ul>\n<p>Unfortunately, we were not able to fit our solution into one single notebook. The reason is that a notebook session on Kaggle is limited to 12 hours. Training one of the seven lightgbm models with our custom loss takes ~11 hours (a notebook in which this training happened has been made public: <a href=\"https://www.kaggle.com/code/gromml/custom-loss\" target=\"_blank\">https://www.kaggle.com/code/gromml/custom-loss</a>. The outputs of the versions 28, 29, 30, 32, 34, 35, 36 were downloaded and published as a public dataset <a href=\"https://www.kaggle.com/datasets/gromml/home-credit-lgb7-custom-loss)\" target=\"_blank\">https://www.kaggle.com/datasets/gromml/home-credit-lgb7-custom-loss)</a>.</p>\n<p>Seven lightgbm models with a default loss were also pre-trained (<a href=\"https://www.kaggle.com/code/gromml/log-loss\" target=\"_blank\">https://www.kaggle.com/code/gromml/log-loss</a>, see the Version 47 output), and another public dataset was created: <a href=\"https://www.kaggle.com/datasets/gromml/home-credit-lgb7\" target=\"_blank\">https://www.kaggle.com/datasets/gromml/home-credit-lgb7</a>.</p>\n<p>Seven catboost models with a default loss were also pre-trained (<a href=\"https://www.kaggle.com/code/gromml/log-loss\" target=\"_blank\">https://www.kaggle.com/code/gromml/log-loss</a>, see the Version 42 output), and another public dataset was created: <a href=\"https://www.kaggle.com/datasets/gromml/home-credit-cat7\" target=\"_blank\">https://www.kaggle.com/datasets/gromml/home-credit-cat7</a>.</p>\n<p>Our final notebook <a href=\"https://www.kaggle.com/code/gromml/custom-loss-stablerocaucloss\" target=\"_blank\">https://www.kaggle.com/code/gromml/custom-loss-stablerocaucloss</a> is considered as our final notebook, where all our results were collected together to produce a competition submission.</p>",
      "rawMarkdown": "The competition metric is given by the following equation:\nstability metric = mean(gini) + 88 * min(0, a) - 0.5 * std(residuals).\n\nAt the beginning of the competition, we wanted to construct a loss consisting of three terms: the first term should have addressed mean(gini), the second one could have been responsible for the time decay, and the last one should have addressed std of residuals. However, when `WEEK_NUM` was replaced with a constant, we decided that we could hardly incorporate decay (that is, the 2nd term) into our loss function without information about time and dates. That is why, we decided to focus on only two terms with the hope of minimizing the decay effect throughout the last term. Indeed, if there is no trend in gini scores, there is a high chance that std of gini scores is lower in comparison with a linear trend in gini scores. This is how we came to our custom loss described in our public notebook: https://www.kaggle.com/code/gromml/custom-loss-stablerocaucloss. In this notebook one can find additional details about our solution (preprocessing, ensembling, etc.).\n\nAs mentioned in this public notebook, our ensemble consists of 21 models:\n* 7 lightgbm models (one model for each of the 7 folds) trained using a default loss;\n* 7 catboost models trained using a default loss (one model for each fold);\n* 7 lightgbm models trained using the custom loss (which was introduced in the public notebook).\n\nUnfortunately, we were not able to fit our solution into one single notebook. The reason is that a notebook session on Kaggle is limited to 12 hours. Training one of the seven lightgbm models with our custom loss takes ~11 hours (a notebook in which this training happened has been made public: https://www.kaggle.com/code/gromml/custom-loss. The outputs of the versions 28, 29, 30, 32, 34, 35, 36 were downloaded and published as a public dataset https://www.kaggle.com/datasets/gromml/home-credit-lgb7-custom-loss).\n\nSeven lightgbm models with a default loss were also pre-trained (https://www.kaggle.com/code/gromml/log-loss, see the Version 47 output), and another public dataset was created: https://www.kaggle.com/datasets/gromml/home-credit-lgb7.\n\nSeven catboost models with a default loss were also pre-trained (https://www.kaggle.com/code/gromml/log-loss, see the Version 42 output), and another public dataset was created: https://www.kaggle.com/datasets/gromml/home-credit-cat7.\n\nOur final notebook https://www.kaggle.com/code/gromml/custom-loss-stablerocaucloss is considered as our final notebook, where all our results were collected together to produce a competition submission.\n",
      "votes": 1
    },
    {
      "id": 2840173,
      "postDate": "2024-05-28T02:14:21.777Z",
      "content": "<p>Excellent work! I am so inspired with the work you have put in. I tried aggregating the similar columns from different table, but I did not see it perform better than filtering based on correlation. But, of course, I did not try averaging. I used max for most of them. Your custom metric function is just exceptional. Thank you so much for sharing your work.</p>",
      "rawMarkdown": "Excellent work! I am so inspired with the work you have put in. I tried aggregating the similar columns from different table, but I did not see it perform better than filtering based on correlation. But, of course, I did not try averaging. I used max for most of them. Your custom metric function is just exceptional. Thank you so much for sharing your work.",
      "votes": 1
    },
    {
      "id": 2840801,
      "postDate": "2024-05-28T09:05:54.820Z",
      "content": "<p>Excellent work!  I am interested in the custom loss. Why not change the feval corresponding to the custom loss.  I found that  the 10 sample predict probability of custom loss(047~0.52) is far bigger than other 2 models(&lt;0.22).  I guess it's performance may vary greatly. Can you disclose the scores without custome loss model or only custome loss model? Thank you!</p>",
      "rawMarkdown": "Excellent work!  I am interested in the custom loss. Why not change the feval corresponding to the custom loss.  I found that  the 10 sample predict probability of custom loss(047~0.52) is far bigger than other 2 models(<0.22).  I guess it's performance may vary greatly. Can you disclose the scores without custome loss model or only custome loss model? Thank you!",
      "replies": [
        {
          "id": 2840877,
          "postDate": "2024-05-28T09:35:13.477Z",
          "content": "<p>When you use a custom loss with a lightgbm classifier, its output is raw scores rather than probabilities. We used the sigmoid function to transform raw scores (~ from -0.3 to 0.3) into [0; 1] values. After that, we average these values with probabilities from classic lightgbm and catboost models.</p>\n<p>When your metric is AUC, the mean value of your predictions is not important. What is important is their standard deviation, and it is approximately the same for custom loss predictions and standard loss predictions. So, different mean values do not make ensemble performance unstable.</p>\n<p>It is also worth mentioning that your final predictions do not have to be in the [0; 1] interval. If your metric is AUC, it is not compulsory to normalize your predictions into [0; 1] interval.</p>",
          "rawMarkdown": "When you use a custom loss with a lightgbm classifier, its output is raw scores rather than probabilities. We used the sigmoid function to transform raw scores (~ from -0.3 to 0.3) into [0; 1] values. After that, we average these values with probabilities from classic lightgbm and catboost models.\n\nWhen your metric is AUC, the mean value of your predictions is not important. What is important is their standard deviation, and it is approximately the same for custom loss predictions and standard loss predictions. So, different mean values do not make ensemble performance unstable.\n\nIt is also worth mentioning that your final predictions do not have to be in the [0; 1] interval. If your metric is AUC, it is not compulsory to normalize your predictions into [0; 1] interval.",
          "votes": 1
        },
        {
          "id": 2841823,
          "postDate": "2024-05-28T17:46:47.133Z",
          "content": "<p>Also, our score without custom loss (that is, standard lightgbm + standard catboost) was 0.582 / 0.502. Adding the lightgbm with the custom loss to our ensemble makes it to be 0.585 / 0.502.</p>\n<p>We didn't submit models trained only with custom loss. You can try submitting only custom loss models and report the results :)</p>",
          "rawMarkdown": "Also, our score without custom loss (that is, standard lightgbm + standard catboost) was 0.582 / 0.502. Adding the lightgbm with the custom loss to our ensemble makes it to be 0.585 / 0.502.\n\nWe didn't submit models trained only with custom loss. You can try submitting only custom loss models and report the results :)\n\n",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2840159,
      "author_name": "Simo Elm",
      "author_url": "",
      "post_date": "2024-05-28T02:02:08.193000",
      "content": "<p>Good and interesting approach. But, you should specify that the .544 is with metric hacking.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2840551,
          "author_name": "Andrey Nesterov",
          "author_url": "",
          "post_date": "2024-05-28T06:20:04.750000",
          "content": "<p>Yes, interesting approach. I've tried it without metric hack, and it showed not excellent but pretty good results: Public - 0.585, Private - 0.502.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2840744,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-05-28T08:43:44.273000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2840738,
          "author_name": "gromml",
          "author_url": "",
          "post_date": "2024-05-28T08:40:12.573000",
          "content": "<p>Yes, 0.544 uses the metric trick. One can check the second version of the notebook, V2 does not utilize any metric tricks. It has 0.585 / 0.502 score.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2869840,
      "author_name": "gromml",
      "author_url": "",
      "post_date": "2024-06-13T09:57:04.173000",
      "content": "<p>Also, there is a file below with some tests of our loss.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2866618,
      "author_name": "gromml",
      "author_url": "",
      "post_date": "2024-06-11T12:07:38.893000",
      "content": "<p>The competition metric is given by the following equation:<br>\nstability metric = mean(gini) + 88 * min(0, a) - 0.5 * std(residuals).</p>\n<p>At the beginning of the competition, we wanted to construct a loss consisting of three terms: the first term should have addressed mean(gini), the second one could have been responsible for the time decay, and the last one should have addressed std of residuals. However, when <code>WEEK_NUM</code> was replaced with a constant, we decided that we could hardly incorporate decay (that is, the 2nd term) into our loss function without information about time and dates. That is why, we decided to focus on only two terms with the hope of minimizing the decay effect throughout the last term. Indeed, if there is no trend in gini scores, there is a high chance that std of gini scores is lower in comparison with a linear trend in gini scores. This is how we came to our custom loss described in our public notebook: <a href=\"https://www.kaggle.com/code/gromml/custom-loss-stablerocaucloss\" target=\"_blank\">https://www.kaggle.com/code/gromml/custom-loss-stablerocaucloss</a>. In this notebook one can find additional details about our solution (preprocessing, ensembling, etc.).</p>\n<p>As mentioned in this public notebook, our ensemble consists of 21 models:</p>\n<ul>\n<li>7 lightgbm models (one model for each of the 7 folds) trained using a default loss;</li>\n<li>7 catboost models trained using a default loss (one model for each fold);</li>\n<li>7 lightgbm models trained using the custom loss (which was introduced in the public notebook).</li>\n</ul>\n<p>Unfortunately, we were not able to fit our solution into one single notebook. The reason is that a notebook session on Kaggle is limited to 12 hours. Training one of the seven lightgbm models with our custom loss takes ~11 hours (a notebook in which this training happened has been made public: <a href=\"https://www.kaggle.com/code/gromml/custom-loss\" target=\"_blank\">https://www.kaggle.com/code/gromml/custom-loss</a>. The outputs of the versions 28, 29, 30, 32, 34, 35, 36 were downloaded and published as a public dataset <a href=\"https://www.kaggle.com/datasets/gromml/home-credit-lgb7-custom-loss)\" target=\"_blank\">https://www.kaggle.com/datasets/gromml/home-credit-lgb7-custom-loss)</a>.</p>\n<p>Seven lightgbm models with a default loss were also pre-trained (<a href=\"https://www.kaggle.com/code/gromml/log-loss\" target=\"_blank\">https://www.kaggle.com/code/gromml/log-loss</a>, see the Version 47 output), and another public dataset was created: <a href=\"https://www.kaggle.com/datasets/gromml/home-credit-lgb7\" target=\"_blank\">https://www.kaggle.com/datasets/gromml/home-credit-lgb7</a>.</p>\n<p>Seven catboost models with a default loss were also pre-trained (<a href=\"https://www.kaggle.com/code/gromml/log-loss\" target=\"_blank\">https://www.kaggle.com/code/gromml/log-loss</a>, see the Version 42 output), and another public dataset was created: <a href=\"https://www.kaggle.com/datasets/gromml/home-credit-cat7\" target=\"_blank\">https://www.kaggle.com/datasets/gromml/home-credit-cat7</a>.</p>\n<p>Our final notebook <a href=\"https://www.kaggle.com/code/gromml/custom-loss-stablerocaucloss\" target=\"_blank\">https://www.kaggle.com/code/gromml/custom-loss-stablerocaucloss</a> is considered as our final notebook, where all our results were collected together to produce a competition submission.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2840173,
      "author_name": "Varuni Rao",
      "author_url": "",
      "post_date": "2024-05-28T02:14:21.777000",
      "content": "<p>Excellent work! I am so inspired with the work you have put in. I tried aggregating the similar columns from different table, but I did not see it perform better than filtering based on correlation. But, of course, I did not try averaging. I used max for most of them. Your custom metric function is just exceptional. Thank you so much for sharing your work.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2840801,
      "author_name": "sequoia12",
      "author_url": "",
      "post_date": "2024-05-28T09:05:54.820000",
      "content": "<p>Excellent work!  I am interested in the custom loss. Why not change the feval corresponding to the custom loss.  I found that  the 10 sample predict probability of custom loss(047~0.52) is far bigger than other 2 models(&lt;0.22).  I guess it's performance may vary greatly. Can you disclose the scores without custome loss model or only custome loss model? Thank you!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2840877,
          "author_name": "gromml",
          "author_url": "",
          "post_date": "2024-05-28T09:35:13.477000",
          "content": "<p>When you use a custom loss with a lightgbm classifier, its output is raw scores rather than probabilities. We used the sigmoid function to transform raw scores (~ from -0.3 to 0.3) into [0; 1] values. After that, we average these values with probabilities from classic lightgbm and catboost models.</p>\n<p>When your metric is AUC, the mean value of your predictions is not important. What is important is their standard deviation, and it is approximately the same for custom loss predictions and standard loss predictions. So, different mean values do not make ensemble performance unstable.</p>\n<p>It is also worth mentioning that your final predictions do not have to be in the [0; 1] interval. If your metric is AUC, it is not compulsory to normalize your predictions into [0; 1] interval.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2841823,
          "author_name": "gromml",
          "author_url": "",
          "post_date": "2024-05-28T17:46:47.133000",
          "content": "<p>Also, our score without custom loss (that is, standard lightgbm + standard catboost) was 0.582 / 0.502. Adding the lightgbm with the custom loss to our ensemble makes it to be 0.585 / 0.502.</p>\n<p>We didn't submit models trained only with custom loss. You can try submitting only custom loss models and report the results :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2840076": "Here is our work with the custom loss (called StableRocAucLoss):\n\nhttps://www.kaggle.com/code/gromml/custom-loss-stablerocaucloss\n\nWill be interesting to see approaches of other participants to address stability!",
    "2840159": "Good and interesting approach. But, you should specify that the .544 is with metric hacking.",
    "2869840": "Also, there is a file below with some tests of our loss.",
    "2866618": "The competition metric is given by the following equation:\nstability metric = mean(gini) + 88 * min(0, a) - 0.5 * std(residuals).\n\nAt the beginning of the competition, we wanted to construct a loss consisting of three terms: the first term should have addressed mean(gini), the second one could have been responsible for the time decay, and the last one should have addressed std of residuals. However, when `WEEK_NUM` was replaced with a constant, we decided that we could hardly incorporate decay (that is, the 2nd term) into our loss function without information about time and dates. That is why, we decided to focus on only two terms with the hope of minimizing the decay effect throughout the last term. Indeed, if there is no trend in gini scores, there is a high chance that std of gini scores is lower in comparison with a linear trend in gini scores. This is how we came to our custom loss described in our public notebook: https://www.kaggle.com/code/gromml/custom-loss-stablerocaucloss. In this notebook one can find additional details about our solution (preprocessing, ensembling, etc.).\n\nAs mentioned in this public notebook, our ensemble consists of 21 models:\n* 7 lightgbm models (one model for each of the 7 folds) trained using a default loss;\n* 7 catboost models trained using a default loss (one model for each fold);\n* 7 lightgbm models trained using the custom loss (which was introduced in the public notebook).\n\nUnfortunately, we were not able to fit our solution into one single notebook. The reason is that a notebook session on Kaggle is limited to 12 hours. Training one of the seven lightgbm models with our custom loss takes ~11 hours (a notebook in which this training happened has been made public: https://www.kaggle.com/code/gromml/custom-loss. The outputs of the versions 28, 29, 30, 32, 34, 35, 36 were downloaded and published as a public dataset https://www.kaggle.com/datasets/gromml/home-credit-lgb7-custom-loss).\n\nSeven lightgbm models with a default loss were also pre-trained (https://www.kaggle.com/code/gromml/log-loss, see the Version 47 output), and another public dataset was created: https://www.kaggle.com/datasets/gromml/home-credit-lgb7.\n\nSeven catboost models with a default loss were also pre-trained (https://www.kaggle.com/code/gromml/log-loss, see the Version 42 output), and another public dataset was created: https://www.kaggle.com/datasets/gromml/home-credit-cat7.\n\nOur final notebook https://www.kaggle.com/code/gromml/custom-loss-stablerocaucloss is considered as our final notebook, where all our results were collected together to produce a competition submission.\n",
    "2840173": "Excellent work! I am so inspired with the work you have put in. I tried aggregating the similar columns from different table, but I did not see it perform better than filtering based on correlation. But, of course, I did not try averaging. I used max for most of them. Your custom metric function is just exceptional. Thank you so much for sharing your work.",
    "2840801": "Excellent work!  I am interested in the custom loss. Why not change the feval corresponding to the custom loss.  I found that  the 10 sample predict probability of custom loss(047~0.52) is far bigger than other 2 models(<0.22).  I guess it's performance may vary greatly. Can you disclose the scores without custome loss model or only custome loss model? Thank you!"
  }
}