{
  "id": 507869,
  "title": "What's the Real Goal? A Few Reflections",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/507869",
  "author_name": "Wagner Assis",
  "post_date": "2024-05-27T15:14:18.806000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>It’s a great competition, and I thank the host for providing and sharing this rich dataset. However, I have some constructive criticism to offer. The dataset provided is comprehensive and closely resembles real-life problems in this field, but I wonder, what is the real goal? In my opinion, the competition seems more like a playground for using real data to capture insights rather than focusing on finding the best model/approach.</p>\n<p>Despite participants’ complaints, the real issue lies in the lack of a specific goal or constraints. None of the showcased solutions with good metric can go into production. We're dealing with the financial sector, one of the most regulated industries. Therefore, the data and the showcased solutions have several shortcomings:<br>\n•    <strong>Lack of <code>model fairness</code> and lack of legal/regulatory compliance</strong> (e.g., the use of the “gender” feature);<br>\n•    <strong>Lack of explainability</strong>: In this field, it is essential to explain the model's decisions for auditing purposes and/or to applicants (in many countries, it’s an applicant’s right). Hence, black-box models or ensembles are not welcome in this field (this is why many companies still use Logistic Regression);<br>\n•    <strong>Lack of flexibility</strong>: There are no time-dependent raw features to work with, such as pmt over time. It would be beneficial to create trends or time-window features beyond the ones given.</p>\n<p>Additionally, application/behavior scores are not developed to use classes, but probabilities/scores to be flexible to policies and risk appetite. Therefore, the competition should focus more on the distribution of probabilities rather than just using <code>.clip()</code>. A good model is not necessarily a confident one, but one that shows stable distribution ratings over time (i.e., ceteris paribus: if there is no major drift, the volumetry and PD by rating/quantiles should remain stable over time).</p>\n<p>For future competitions, it would be great to invest more time in defining what constitutes a “well-done” solution rather than just focusing on metrics. What’s the point of rewarding a black-box model, in this example, if it can never go into production due compliance or regulatory matters? Beyond metrics, I would be good if you would have a discretionary budget to reward the best ideas, processes, and insights. I believe this would encourage participants to showcase and share real solutions instead of merely engaging in metric hacking.</p>\n<p>Regards,</p>",
  "messages": [
    {
      "id": 2839539,
      "postDate": "2024-05-27T15:14:18.807Z",
      "content": "<p>It’s a great competition, and I thank the host for providing and sharing this rich dataset. However, I have some constructive criticism to offer. The dataset provided is comprehensive and closely resembles real-life problems in this field, but I wonder, what is the real goal? In my opinion, the competition seems more like a playground for using real data to capture insights rather than focusing on finding the best model/approach.</p>\n<p>Despite participants’ complaints, the real issue lies in the lack of a specific goal or constraints. None of the showcased solutions with good metric can go into production. We're dealing with the financial sector, one of the most regulated industries. Therefore, the data and the showcased solutions have several shortcomings:<br>\n•    <strong>Lack of <code>model fairness</code> and lack of legal/regulatory compliance</strong> (e.g., the use of the “gender” feature);<br>\n•    <strong>Lack of explainability</strong>: In this field, it is essential to explain the model's decisions for auditing purposes and/or to applicants (in many countries, it’s an applicant’s right). Hence, black-box models or ensembles are not welcome in this field (this is why many companies still use Logistic Regression);<br>\n•    <strong>Lack of flexibility</strong>: There are no time-dependent raw features to work with, such as pmt over time. It would be beneficial to create trends or time-window features beyond the ones given.</p>\n<p>Additionally, application/behavior scores are not developed to use classes, but probabilities/scores to be flexible to policies and risk appetite. Therefore, the competition should focus more on the distribution of probabilities rather than just using <code>.clip()</code>. A good model is not necessarily a confident one, but one that shows stable distribution ratings over time (i.e., ceteris paribus: if there is no major drift, the volumetry and PD by rating/quantiles should remain stable over time).</p>\n<p>For future competitions, it would be great to invest more time in defining what constitutes a “well-done” solution rather than just focusing on metrics. What’s the point of rewarding a black-box model, in this example, if it can never go into production due compliance or regulatory matters? Beyond metrics, I would be good if you would have a discretionary budget to reward the best ideas, processes, and insights. I believe this would encourage participants to showcase and share real solutions instead of merely engaging in metric hacking.</p>\n<p>Regards,</p>",
      "rawMarkdown": "It’s a great competition, and I thank the host for providing and sharing this rich dataset. However, I have some constructive criticism to offer. The dataset provided is comprehensive and closely resembles real-life problems in this field, but I wonder, what is the real goal? In my opinion, the competition seems more like a playground for using real data to capture insights rather than focusing on finding the best model/approach.\n\nDespite participants’ complaints, the real issue lies in the lack of a specific goal or constraints. None of the showcased solutions with good metric can go into production. We're dealing with the financial sector, one of the most regulated industries. Therefore, the data and the showcased solutions have several shortcomings:\n•\t**Lack of `model fairness` and lack of legal/regulatory compliance** (e.g., the use of the “gender” feature);\n•\t**Lack of explainability**: In this field, it is essential to explain the model's decisions for auditing purposes and/or to applicants (in many countries, it’s an applicant’s right). Hence, black-box models or ensembles are not welcome in this field (this is why many companies still use Logistic Regression);\n•\t**Lack of flexibility**: There are no time-dependent raw features to work with, such as pmt over time. It would be beneficial to create trends or time-window features beyond the ones given.\n\nAdditionally, application/behavior scores are not developed to use classes, but probabilities/scores to be flexible to policies and risk appetite. Therefore, the competition should focus more on the distribution of probabilities rather than just using `.clip()`. A good model is not necessarily a confident one, but one that shows stable distribution ratings over time (i.e., ceteris paribus: if there is no major drift, the volumetry and PD by rating/quantiles should remain stable over time).\n\nFor future competitions, it would be great to invest more time in defining what constitutes a “well-done” solution rather than just focusing on metrics. What’s the point of rewarding a black-box model, in this example, if it can never go into production due compliance or regulatory matters? Beyond metrics, I would be good if you would have a discretionary budget to reward the best ideas, processes, and insights. I believe this would encourage participants to showcase and share real solutions instead of merely engaging in metric hacking.\n\nRegards,",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2839539": "It’s a great competition, and I thank the host for providing and sharing this rich dataset. However, I have some constructive criticism to offer. The dataset provided is comprehensive and closely resembles real-life problems in this field, but I wonder, what is the real goal? In my opinion, the competition seems more like a playground for using real data to capture insights rather than focusing on finding the best model/approach.\n\nDespite participants’ complaints, the real issue lies in the lack of a specific goal or constraints. None of the showcased solutions with good metric can go into production. We're dealing with the financial sector, one of the most regulated industries. Therefore, the data and the showcased solutions have several shortcomings:\n•\t**Lack of `model fairness` and lack of legal/regulatory compliance** (e.g., the use of the “gender” feature);\n•\t**Lack of explainability**: In this field, it is essential to explain the model's decisions for auditing purposes and/or to applicants (in many countries, it’s an applicant’s right). Hence, black-box models or ensembles are not welcome in this field (this is why many companies still use Logistic Regression);\n•\t**Lack of flexibility**: There are no time-dependent raw features to work with, such as pmt over time. It would be beneficial to create trends or time-window features beyond the ones given.\n\nAdditionally, application/behavior scores are not developed to use classes, but probabilities/scores to be flexible to policies and risk appetite. Therefore, the competition should focus more on the distribution of probabilities rather than just using `.clip()`. A good model is not necessarily a confident one, but one that shows stable distribution ratings over time (i.e., ceteris paribus: if there is no major drift, the volumetry and PD by rating/quantiles should remain stable over time).\n\nFor future competitions, it would be great to invest more time in defining what constitutes a “well-done” solution rather than just focusing on metrics. What’s the point of rewarding a black-box model, in this example, if it can never go into production due compliance or regulatory matters? Beyond metrics, I would be good if you would have a discretionary budget to reward the best ideas, processes, and insights. I believe this would encourage participants to showcase and share real solutions instead of merely engaging in metric hacking.\n\nRegards,"
  }
}