{
  "id": 334157,
  "title": "G and D learning curves",
  "url": "/competitions/amex-default-prediction/discussion/334157",
  "author_name": "",
  "post_date": "2022-06-30T04:42:56.869210Z",
  "votes": 30,
  "comment_count": 6,
  "views": 0,
  "content": "<p>As explained by the organizers <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/overview/evaluation\" target=\"_blank\">here</a>, the competition's metric <strong>M is the average of two sub-metrics: G and D</strong>.<br>\n In this discussion, we will look at the learning curve for each sub-metric, while training an XGB model. During the cross validation process, a validation subset of the training data is kept apart (not shown to the model).  The curves below show the evolution of G and D on this validation subset, during the XGB training. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F979b3f044b3a7278393822fcdfdefd88%2FGandDLearningCurves.png?generation=1656602775384215&amp;alt=media\" alt=\"\"></p>\n<p>The model converges at an M score of 0.793 but the sub-metrics have very different behaviors. G reaches 0.924 while D merely reaches 0.662. This seems to indicate that there is more room for improvement on the D sub-metric. Unfortunately, we only see M on the LeaderBoard.</p>\n<p><a href=\"https://www.kaggle.com/code/gehallak/the-dark-side-of-the-moon\" target=\"_blank\">This Notebook</a> looks into D in more details.</p>",
  "messages": [
    {
      "id": "1837967",
      "postDate": "06/30/2022 04:42:56",
      "content": "<p>As explained by the organizers <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/overview/evaluation\" target=\"_blank\">here</a>, the competition's metric <strong>M is the average of two sub-metrics: G and D</strong>.<br>\n In this discussion, we will look at the learning curve for each sub-metric, while training an XGB model. During the cross validation process, a validation subset of the training data is kept apart (not shown to the model).  The curves below show the evolution of G and D on this validation subset, during the XGB training. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F979b3f044b3a7278393822fcdfdefd88%2FGandDLearningCurves.png?generation=1656602775384215&amp;alt=media\" alt=\"\"></p>\n<p>The model converges at an M score of 0.793 but the sub-metrics have very different behaviors. G reaches 0.924 while D merely reaches 0.662. This seems to indicate that there is more room for improvement on the D sub-metric. Unfortunately, we only see M on the LeaderBoard.</p>\n<p><a href=\"https://www.kaggle.com/code/gehallak/the-dark-side-of-the-moon\" target=\"_blank\">This Notebook</a> looks into D in more details.</p>",
      "rawMarkdown": "As explained by the organizers [here](https://www.kaggle.com/competitions/amex-default-prediction/overview/evaluation), the competition's metric **M is the average of two sub-metrics: G and D**.\n In this discussion, we will look at the learning curve for each sub-metric, while training an XGB model. During the cross validation process, a validation subset of the training data is kept apart (not shown to the model).  The curves below show the evolution of G and D on this validation subset, during the XGB training. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F979b3f044b3a7278393822fcdfdefd88%2FGandDLearningCurves.png?generation=1656602775384215&alt=media)\n\nThe model converges at an M score of 0.793 but the sub-metrics have very different behaviors. G reaches 0.924 while D merely reaches 0.662. This seems to indicate that there is more room for improvement on the D sub-metric. Unfortunately, we only see M on the LeaderBoard.\n\n[This Notebook](https://www.kaggle.com/code/gehallak/the-dark-side-of-the-moon) looks into D in more details.",
      "votes": null
    },
    {
      "id": "1838764",
      "postDate": "06/30/2022 19:27:29",
      "content": "<p>So, there’s a strong bias toward high precision at the given recall. I wondered how the objective function could be modified to take advantage of this. Cue diversion into focal loss and weighted loss. I come back empty handed, but somewhat wiser. My next idea. Look at what errors my model is making in the competition metric. Is it for a certain type of customer?</p>\n<p>Ps you could also plot logloss, the actual objective function.</p>",
      "rawMarkdown": "So, there’s a strong bias toward high precision at the given recall. I wondered how the objective function could be modified to take advantage of this. Cue diversion into focal loss and weighted loss. I come back empty handed, but somewhat wiser. My next idea. Look at what errors my model is making in the competition metric. Is it for a certain type of customer?\n\nPs you could also plot logloss, the actual objective function.",
      "votes": null
    },
    {
      "id": "1839931",
      "postDate": "07/01/2022 20:03:05",
      "content": "<p>I like the idea of the possibility to squeeze some more of D by giving away some of G.</p>\n<p>There is <a href=\"https://arxiv.org/pdf/2009.14119.pdf\" target=\"_blank\">this paper</a> and the related <a href=\"https://github.com/Alibaba-MIIL/ASL\" target=\"_blank\">github repo</a> that could be helpful. They essentially implement a modified (asymmetric) focal cross-entropy loss by applying different focal strengths for the hard false positives and hard false negatives, as opposed to the standard focal cross-entropy where hard to classify samples are weighted equally for positive and negative labels.</p>\n<p>In my case it did not improve the performance though. It seems that G and D are too tightly coupled. Maybe some else has better luck with it.</p>",
      "rawMarkdown": "I like the idea of the possibility to squeeze some more of D by giving away some of G.\n\nThere is [this paper](https://arxiv.org/pdf/2009.14119.pdf) and the related [github repo](https://github.com/Alibaba-MIIL/ASL) that could be helpful. They essentially implement a modified (asymmetric) focal cross-entropy loss by applying different focal strengths for the hard false positives and hard false negatives, as opposed to the standard focal cross-entropy where hard to classify samples are weighted equally for positive and negative labels.\n\nIn my case it did not improve the performance though. It seems that G and D are too tightly coupled. Maybe some else has better luck with it.",
      "votes": null
    },
    {
      "id": "1840244",
      "postDate": "07/02/2022 04:38:39",
      "content": "<p>it's an interesting idea to try different loss functions. This is actually why I wanted to see these curves. It appeared clearly when drawing these curves, that during the XGB  learning process, while the logloss goes down regularly and smoothly,  G and D go up regularly and smoothly too. The 3 of them converge to their asymptotes together without erratic moves. Empirically, it looks like even if the model knows nothing about the metrics G and D, it is doing a very good job at maximizing them. This is probably because we have a large number of customers. Trying to accurately predict the probability of default for each of them translates correctly in ranking them in order of probability of default (which is the only thing G and D care about).</p>",
      "rawMarkdown": "it's an interesting idea to try different loss functions. This is actually why I wanted to see these curves. It appeared clearly when drawing these curves, that during the XGB  learning process, while the logloss goes down regularly and smoothly,  G and D go up regularly and smoothly too. The 3 of them converge to their asymptotes together without erratic moves. Empirically, it looks like even if the model knows nothing about the metrics G and D, it is doing a very good job at maximizing them. This is probably because we have a large number of customers. Trying to accurately predict the probability of default for each of them translates correctly in ranking them in order of probability of default (which is the only thing G and D care about).",
      "votes": null
    },
    {
      "id": "1844191",
      "postDate": "07/05/2022 12:04:22",
      "content": "<p>It takes real courage to land on the dark side of the moon, thanks for sharing your knowledge with us!</p>",
      "rawMarkdown": "It takes real courage to land on the dark side of the moon, thanks for sharing your knowledge with us!",
      "votes": null
    },
    {
      "id": "1844341",
      "postDate": "07/05/2022 13:34:27",
      "content": "<p>Thanks for reading and for your sense of humor 😂</p>",
      "rawMarkdown": "Thanks for reading and for your sense of humor 😂",
      "votes": null
    },
    {
      "id": "1844350",
      "postDate": "07/05/2022 13:46:54",
      "content": "<p>Thanks for your thoughts <a href=\"https://www.kaggle.com/burritodan\" target=\"_blank\">@burritodan</a>. I have tried to focus on ranking (which is all the metrics care about) instead of prediction (loss function) during the blending but without any noticeable increase (or decrease) in the LB. As if with a large number of customers, both problems converge.</p>",
      "rawMarkdown": "Thanks for your thoughts @burritodan. I have tried to focus on ranking (which is all the metrics care about) instead of prediction (loss function) during the blending but without any noticeable increase (or decrease) in the LB. As if with a large number of customers, both problems converge.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1838764,
      "author_name": "burritodan",
      "author_url": "",
      "post_date": "06/30/2022 19:27:29",
      "content": "<p>So, there’s a strong bias toward high precision at the given recall. I wondered how the objective function could be modified to take advantage of this. Cue diversion into focal loss and weighted loss. I come back empty handed, but somewhat wiser. My next idea. Look at what errors my model is making in the competition metric. Is it for a certain type of customer?</p>\n<p>Ps you could also plot logloss, the actual objective function.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1844350,
          "author_name": "gehallak",
          "author_url": "",
          "post_date": "07/05/2022 13:46:54",
          "content": "<p>Thanks for your thoughts <a href=\"https://www.kaggle.com/burritodan\" target=\"_blank\">@burritodan</a>. I have tried to focus on ranking (which is all the metrics care about) instead of prediction (loss function) during the blending but without any noticeable increase (or decrease) in the LB. As if with a large number of customers, both problems converge.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1839931,
      "author_name": "kaerim",
      "author_url": "",
      "post_date": "07/01/2022 20:03:05",
      "content": "<p>I like the idea of the possibility to squeeze some more of D by giving away some of G.</p>\n<p>There is <a href=\"https://arxiv.org/pdf/2009.14119.pdf\" target=\"_blank\">this paper</a> and the related <a href=\"https://github.com/Alibaba-MIIL/ASL\" target=\"_blank\">github repo</a> that could be helpful. They essentially implement a modified (asymmetric) focal cross-entropy loss by applying different focal strengths for the hard false positives and hard false negatives, as opposed to the standard focal cross-entropy where hard to classify samples are weighted equally for positive and negative labels.</p>\n<p>In my case it did not improve the performance though. It seems that G and D are too tightly coupled. Maybe some else has better luck with it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1840244,
          "author_name": "gehallak",
          "author_url": "",
          "post_date": "07/02/2022 04:38:39",
          "content": "<p>it's an interesting idea to try different loss functions. This is actually why I wanted to see these curves. It appeared clearly when drawing these curves, that during the XGB  learning process, while the logloss goes down regularly and smoothly,  G and D go up regularly and smoothly too. The 3 of them converge to their asymptotes together without erratic moves. Empirically, it looks like even if the model knows nothing about the metrics G and D, it is doing a very good job at maximizing them. This is probably because we have a large number of customers. Trying to accurately predict the probability of default for each of them translates correctly in ranking them in order of probability of default (which is the only thing G and D care about).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1844191,
      "author_name": "saurabhbagchi",
      "author_url": "",
      "post_date": "07/05/2022 12:04:22",
      "content": "<p>It takes real courage to land on the dark side of the moon, thanks for sharing your knowledge with us!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1844341,
          "author_name": "gehallak",
          "author_url": "",
          "post_date": "07/05/2022 13:34:27",
          "content": "<p>Thanks for reading and for your sense of humor 😂</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1837967": "As explained by the organizers [here](https://www.kaggle.com/competitions/amex-default-prediction/overview/evaluation), the competition's metric **M is the average of two sub-metrics: G and D**.\n In this discussion, we will look at the learning curve for each sub-metric, while training an XGB model. During the cross validation process, a validation subset of the training data is kept apart (not shown to the model).  The curves below show the evolution of G and D on this validation subset, during the XGB training. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F979b3f044b3a7278393822fcdfdefd88%2FGandDLearningCurves.png?generation=1656602775384215&alt=media)\n\nThe model converges at an M score of 0.793 but the sub-metrics have very different behaviors. G reaches 0.924 while D merely reaches 0.662. This seems to indicate that there is more room for improvement on the D sub-metric. Unfortunately, we only see M on the LeaderBoard.\n\n[This Notebook](https://www.kaggle.com/code/gehallak/the-dark-side-of-the-moon) looks into D in more details.",
    "1838764": "So, there’s a strong bias toward high precision at the given recall. I wondered how the objective function could be modified to take advantage of this. Cue diversion into focal loss and weighted loss. I come back empty handed, but somewhat wiser. My next idea. Look at what errors my model is making in the competition metric. Is it for a certain type of customer?\n\nPs you could also plot logloss, the actual objective function.",
    "1839931": "I like the idea of the possibility to squeeze some more of D by giving away some of G.\n\nThere is [this paper](https://arxiv.org/pdf/2009.14119.pdf) and the related [github repo](https://github.com/Alibaba-MIIL/ASL) that could be helpful. They essentially implement a modified (asymmetric) focal cross-entropy loss by applying different focal strengths for the hard false positives and hard false negatives, as opposed to the standard focal cross-entropy where hard to classify samples are weighted equally for positive and negative labels.\n\nIn my case it did not improve the performance though. It seems that G and D are too tightly coupled. Maybe some else has better luck with it.",
    "1840244": "it's an interesting idea to try different loss functions. This is actually why I wanted to see these curves. It appeared clearly when drawing these curves, that during the XGB  learning process, while the logloss goes down regularly and smoothly,  G and D go up regularly and smoothly too. The 3 of them converge to their asymptotes together without erratic moves. Empirically, it looks like even if the model knows nothing about the metrics G and D, it is doing a very good job at maximizing them. This is probably because we have a large number of customers. Trying to accurately predict the probability of default for each of them translates correctly in ranking them in order of probability of default (which is the only thing G and D care about).",
    "1844191": "It takes real courage to land on the dark side of the moon, thanks for sharing your knowledge with us!",
    "1844341": "Thanks for reading and for your sense of humor 😂",
    "1844350": "Thanks for your thoughts @burritodan. I have tried to focus on ranking (which is all the metrics care about) instead of prediction (loss function) during the blending but without any noticeable increase (or decrease) in the LB. As if with a large number of customers, both problems converge."
  },
  "source": "meta"
}