{
  "id": 483559,
  "title": "Chiming in with a few thoughts",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/483559",
  "author_name": "",
  "post_date": "2024-03-12T23:24:47.099582700Z",
  "votes": 55,
  "comment_count": 16,
  "views": 0,
  "content": "<p>As I've been reading the question, comments, criticisms, and other feedback about the metric and test dataset update, a few things have come to mind that I'd like to share.</p>\n<p>I actively competed in Kaggle competitions for about 5 years before I joined Kaggle 7 years ago. The one thing that really struck me when I was on the \"other side\" was just how difficult it is to launch a successful competition. There are so many things that can go wrong with the data, so many ways leakage can sneak through, and many many other issues that can arise. But even with the (often embarrassing) hiccups that have occurred through the years, I believe the Kaggle platform has had an amazing track record of giving the community fun, interesting, and challenging competitions, year after year.</p>\n<p>We often have to make hard decisions when it comes to launching competitions. There are competitions that come to us in great shape. There are competitions that just aren't feasible that we turn down. And then there are quite a few in the middle, where we have to make \"design decisions\" that might not be optimally tuned for a host or the community. We generally take the stance that it is better to launch an interesting competition with some rough edges than not to launch it at all.</p>\n<p>I believe that Home Credit has provided the community with a very interesting and valuable multi-table dataset, something we don't see very often on Kaggle these days. My hope is that, for those who are interested in this type of problem, you can look past the rough edges and still find your participation valuable. Some will choose not to, which is perfectly fine as well. I personally feel the weight when individuals in the community are frustrated or feel like they've wasted their time in a competition. I sincerely hope you'll stick around, and look forward to the interesting approaches for tackling a tabular problem of this complexity.</p>\n<p>Thanks for listening!</p>",
  "messages": [
    {
      "id": "2694213",
      "postDate": "03/12/2024 23:24:47",
      "content": "<p>As I've been reading the question, comments, criticisms, and other feedback about the metric and test dataset update, a few things have come to mind that I'd like to share.</p>\n<p>I actively competed in Kaggle competitions for about 5 years before I joined Kaggle 7 years ago. The one thing that really struck me when I was on the \"other side\" was just how difficult it is to launch a successful competition. There are so many things that can go wrong with the data, so many ways leakage can sneak through, and many many other issues that can arise. But even with the (often embarrassing) hiccups that have occurred through the years, I believe the Kaggle platform has had an amazing track record of giving the community fun, interesting, and challenging competitions, year after year.</p>\n<p>We often have to make hard decisions when it comes to launching competitions. There are competitions that come to us in great shape. There are competitions that just aren't feasible that we turn down. And then there are quite a few in the middle, where we have to make \"design decisions\" that might not be optimally tuned for a host or the community. We generally take the stance that it is better to launch an interesting competition with some rough edges than not to launch it at all.</p>\n<p>I believe that Home Credit has provided the community with a very interesting and valuable multi-table dataset, something we don't see very often on Kaggle these days. My hope is that, for those who are interested in this type of problem, you can look past the rough edges and still find your participation valuable. Some will choose not to, which is perfectly fine as well. I personally feel the weight when individuals in the community are frustrated or feel like they've wasted their time in a competition. I sincerely hope you'll stick around, and look forward to the interesting approaches for tackling a tabular problem of this complexity.</p>\n<p>Thanks for listening!</p>",
      "rawMarkdown": "As I've been reading the question, comments, criticisms, and other feedback about the metric and test dataset update, a few things have come to mind that I'd like to share.\n\nI actively competed in Kaggle competitions for about 5 years before I joined Kaggle 7 years ago. The one thing that really struck me when I was on the \"other side\" was just how difficult it is to launch a successful competition. There are so many things that can go wrong with the data, so many ways leakage can sneak through, and many many other issues that can arise. But even with the (often embarrassing) hiccups that have occurred through the years, I believe the Kaggle platform has had an amazing track record of giving the community fun, interesting, and challenging competitions, year after year.\n\nWe often have to make hard decisions when it comes to launching competitions. There are competitions that come to us in great shape. There are competitions that just aren't feasible that we turn down. And then there are quite a few in the middle, where we have to make \"design decisions\" that might not be optimally tuned for a host or the community. We generally take the stance that it is better to launch an interesting competition with some rough edges than not to launch it at all.\n\nI believe that Home Credit has provided the community with a very interesting and valuable multi-table dataset, something we don't see very often on Kaggle these days. My hope is that, for those who are interested in this type of problem, you can look past the rough edges and still find your participation valuable. Some will choose not to, which is perfectly fine as well. I personally feel the weight when individuals in the community are frustrated or feel like they've wasted their time in a competition. I sincerely hope you'll stick around, and look forward to the interesting approaches for tackling a tabular problem of this complexity.\n\nThanks for listening!",
      "votes": null
    },
    {
      "id": "2694266",
      "postDate": "03/13/2024 00:19:52",
      "content": "<p>I think the main problem here is confusion about what exactly is going on. For example, all of my old submissions now show errors, however, some seems to have been rescored on the new test data, because they started to show significantly lower scores. At the same time, a few old submissions still show old higher scores that makes my submission list look like a mess. Since we already spent weeks on building our models, it is obvious that this situation doesn't affect motivation in a positive way.</p>",
      "rawMarkdown": "I think the main problem here is confusion about what exactly is going on. For example, all of my old submissions now show errors, however, some seems to have been rescored on the new test data, because they started to show significantly lower scores. At the same time, a few old submissions still show old higher scores that makes my submission list look like a mess. Since we already spent weeks on building our models, it is obvious that this situation doesn't affect motivation in a positive way.",
      "votes": null
    },
    {
      "id": "2694272",
      "postDate": "03/13/2024 00:34:56",
      "content": "<p>All of the previous of the Notebooks that were scored against the old test data were invalidated. We didn't re-run any automatically with the new test data. Did you re-submit these? I'm curious to understand this.</p>",
      "rawMarkdown": "All of the previous of the Notebooks that were scored against the old test data were invalidated. We didn't re-run any automatically with the new test data. Did you re-submit these? I'm curious to understand this.",
      "votes": null
    },
    {
      "id": "2694275",
      "postDate": "03/13/2024 00:46:30",
      "content": "<p>I only resubmitted a few notebooks and found that their scores dropped tremendously, like x2. My guess is that may be the test data have been shuffled and the model somehow expected they're sorted by date or <code>case_id</code>, still trying to figure out the reason.</p>\n<p>All the other submissions are now marked as Error/Succeeded, that I don't know how to interpret:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2099265%2F31ec60d703d57a815069305627cbf90d%2FScreenshot%202024-03-12%20at%205.44.59%20PM.png?generation=1710290723618474&amp;alt=media\"></p>\n<p>If you confirm they were not rescored, then it is easier to simply forget about those LB scores and start from scratch.</p>",
      "rawMarkdown": "I only resubmitted a few notebooks and found that their scores dropped tremendously, like x2. My guess is that may be the test data have been shuffled and the model somehow expected they're sorted by date or `case_id`, still trying to figure out the reason.\n\nAll the other submissions are now marked as Error/Succeeded, that I don't know how to interpret:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2099265%2F31ec60d703d57a815069305627cbf90d%2FScreenshot%202024-03-12%20at%205.44.59%20PM.png?generation=1710290723618474&alt=media)\n\nIf you confirm they were not rescored, then it is easier to simply forget about those LB scores and start from scratch.",
      "votes": null
    },
    {
      "id": "2695022",
      "postDate": "03/13/2024 12:17:02",
      "content": "<p>More competitions is generally better, but when a competition turns into a gamble, it affects not only the competition itself but the value of ranks and medals on the platform. Competition hosts benefit from the presence of competitors who are willing to work for little expected monetary reward in exchange for these platform accolades, so I think it is not unreasonable that hosts be held to some standards of security and robustness. </p>\n<p>Would Kaggle not be willing to consider adopting some objective rules for the metrics it allows to avoid the most transparent hackability risks? The problems with this competition could have been avoided if you required that metrics be monotonic with respect to sample scores.</p>",
      "rawMarkdown": "More competitions is generally better, but when a competition turns into a gamble, it affects not only the competition itself but the value of ranks and medals on the platform. Competition hosts benefit from the presence of competitors who are willing to work for little expected monetary reward in exchange for these platform accolades, so I think it is not unreasonable that hosts be held to some standards of security and robustness. \n\nWould Kaggle not be willing to consider adopting some objective rules for the metrics it allows to avoid the most transparent hackability risks? The problems with this competition could have been avoided if you required that metrics be monotonic with respect to sample scores.",
      "votes": null
    },
    {
      "id": "2695064",
      "postDate": "03/13/2024 12:59:12",
      "content": "<p>I fully understand both the author's and your points. And one of the common rights given to Kagglers is not to participate in competitions they don't like (since they've restarted, there's no going back).</p>",
      "rawMarkdown": "I fully understand both the author's and your points. And one of the common rights given to Kagglers is not to participate in competitions they don't like (since they've restarted, there's no going back).",
      "votes": null
    },
    {
      "id": "2695073",
      "postDate": "03/13/2024 13:12:44",
      "content": "<p>As I just said, when points and medals are distributed randomly it degrades their value. My post is about Kaggle policy going forward, not just this competition.</p>",
      "rawMarkdown": "As I just said, when points and medals are distributed randomly it degrades their value. My post is about Kaggle policy going forward, not just this competition.",
      "votes": null
    },
    {
      "id": "2695088",
      "postDate": "03/13/2024 13:25:39",
      "content": "<p>As the former Kaggler author said, I understand that Kaggle competitions cannot eliminate such imperfections. Also, even if a competition is found to have a flaw, it probably won't end mid-way. I will not be participating in this competition, but I will keep an eye on how it goes.</p>",
      "rawMarkdown": "As the former Kaggler author said, I understand that Kaggle competitions cannot eliminate such imperfections. Also, even if a competition is found to have a flaw, it probably won't end mid-way. I will not be participating in this competition, but I will keep an eye on how it goes.",
      "votes": null
    },
    {
      "id": "2695094",
      "postDate": "03/13/2024 13:29:56",
      "content": "<p>Hackability that arises from metrics that are not monotonic wrt sample scores can be eliminated by requiring that chosen metrics are monotonic wrt sample scores. I did not suggest ending this competition mid-way.</p>",
      "rawMarkdown": "Hackability that arises from metrics that are not monotonic wrt sample scores can be eliminated by requiring that chosen metrics are monotonic wrt sample scores. I did not suggest ending this competition mid-way.",
      "votes": null
    },
    {
      "id": "2695647",
      "postDate": "03/13/2024 19:33:49",
      "content": "<p>I resubmitted a notebook exactly the same and worked <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> <a href=\"https://www.kaggle.com/kononenko\" target=\"_blank\">@kononenko</a> . Other notebooks do not work and need to be rewritten (weekend work for me! :))</p>",
      "rawMarkdown": "I resubmitted a notebook exactly the same and worked @inversion @kononenko . Other notebooks do not work and need to be rewritten (weekend work for me! :))",
      "votes": null
    },
    {
      "id": "2695649",
      "postDate": "03/13/2024 19:35:05",
      "content": "<p>Thanks for posting <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> . In my case, I am agree with you. For me it is an honor to participate in Kaggle competition and I understand if a competition has issues. There are worse things in life, even I also understand that we put a lot of effort in them. We all need to understand and be grateful for the massive amount of data and problems we can work on here. Cheers!</p>",
      "rawMarkdown": "Thanks for posting @inversion . In my case, I am agree with you. For me it is an honor to participate in Kaggle competition and I understand if a competition has issues. There are worse things in life, even I also understand that we put a lot of effort in them. We all need to understand and be grateful for the massive amount of data and problems we can work on here. Cheers!",
      "votes": null
    },
    {
      "id": "2695887",
      "postDate": "03/13/2024 23:26:53",
      "content": "<p>I agree <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> that this competition’s dataset is very unique for Kaggle. It is very realistic in comparison to what data scientists deal with in the “real” tabular world.  Data isn’t always clean and is often messy like this one … as i mentioned <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/482636\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "I agree @inversion that this competition’s dataset is very unique for Kaggle. It is very realistic in comparison to what data scientists deal with in the “real” tabular world.  Data isn’t always clean and is often messy like this one … as i mentioned [here](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/482636)",
      "votes": null
    },
    {
      "id": "2696273",
      "postDate": "03/14/2024 08:04:13",
      "content": "<blockquote>\n  <p>The one thing that really struck me when I was on the \"other side\" was just how difficult it is to launch a successful competition. There are so many things that can go wrong with the data, so many ways leakage can sneak through, and many many other issues that can arise. </p>\n</blockquote>\n<p>You mention only one side of the equation here. The other one is that you have thousands of super-smart, motivated people who are looking at the competition and will point out EVERY error there is. Which is a good thing. </p>",
      "rawMarkdown": ">  The one thing that really struck me when I was on the \"other side\" was just how difficult it is to launch a successful competition. There are so many things that can go wrong with the data, so many ways leakage can sneak through, and many many other issues that can arise. \n\nYou mention only one side of the equation here. The other one is that you have thousands of super-smart, motivated people who are looking at the competition and will point out EVERY error there is. Which is a good thing.",
      "votes": null
    },
    {
      "id": "2697405",
      "postDate": "03/14/2024 21:58:07",
      "content": "<p>I've seen quite a few brilliant researches have flaws in their data discovered by the Kaggle community.</p>",
      "rawMarkdown": "I've seen quite a few brilliant researches have flaws in their data discovered by the Kaggle community.",
      "votes": null
    },
    {
      "id": "2699811",
      "postDate": "03/16/2024 06:45:31",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> , Thank you for your work that is not always simple.</p>",
      "rawMarkdown": "Hi @inversion , Thank you for your work that is not always simple.",
      "votes": null
    },
    {
      "id": "2699850",
      "postDate": "03/16/2024 07:40:03",
      "content": "<p>I am currently in a competition and have seen participants brought these points about the flaws in the given data. </p>",
      "rawMarkdown": "I am currently in a competition and have seen participants brought these points about the flaws in the given data.",
      "votes": null
    },
    {
      "id": "2703270",
      "postDate": "03/18/2024 04:52:37",
      "content": "<p>Yes, I do agree with <a href=\"https://www.kaggle.com/INVERSION\" target=\"_blank\">@INVERSION</a> that the Home credit has provided a multi-table data which we often don't see not only in Kaggle competitions but also the ones that I have seen outside. This is how the real life problems in AI look like and we spend lot of time on data. </p>\n<p>Here is my approach which I would like to share with this community,</p>\n<ol>\n<li><p>The very first step is to understand the data (all the CVS/Parquet files) holistically. For that one needs to understand the relationship between the data in different files. Look at these files like tables in a RDBMS and establish relationship between them. This is very time consuming process since we need to understand the what each column name means and its relevance in solving this problem.</p></li>\n<li><p>Domain Knowledge - This will help in understanding the data and process flow better. There are good videos and text material on Internet. </p></li>\n<li><p>Its good use a drawing board to draw these files (as tables) and try to visually link them.</p></li>\n<li><p>EDA (of course) is an important task and I don't need to tell you experts. </p></li>\n<li><p>Divide the TRAIN data into Train (70%)  and Validate (30%) and apply on different ML models.</p></li>\n<li><p>Publish and see the scores on leader board. </p></li>\n<li><p>Last but not the least, have a team of 2 to 3 people to brainstorm ideas. </p></li>\n</ol>\n<p>Happy Winning !!! </p>",
      "rawMarkdown": "Yes, I do agree with @INVERSION that the Home credit has provided a multi-table data which we often don't see not only in Kaggle competitions but also the ones that I have seen outside. This is how the real life problems in AI look like and we spend lot of time on data. \n\nHere is my approach which I would like to share with this community,\n\n1. The very first step is to understand the data (all the CVS/Parquet files) holistically. For that one needs to understand the relationship between the data in different files. Look at these files like tables in a RDBMS and establish relationship between them. This is very time consuming process since we need to understand the what each column name means and its relevance in solving this problem.\n\n2. Domain Knowledge - This will help in understanding the data and process flow better. There are good videos and text material on Internet. \n\n3. Its good use a drawing board to draw these files (as tables) and try to visually link them.\n\n4. EDA (of course) is an important task and I don't need to tell you experts. \n\n5. Divide the TRAIN data into Train (70%)  and Validate (30%) and apply on different ML models.\n\n6. Publish and see the scores on leader board. \n\n7. Last but not the least, have a team of 2 to 3 people to brainstorm ideas. \n\nHappy Winning !!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2694266,
      "author_name": "kononenko",
      "author_url": "",
      "post_date": "03/13/2024 00:19:52",
      "content": "<p>I think the main problem here is confusion about what exactly is going on. For example, all of my old submissions now show errors, however, some seems to have been rescored on the new test data, because they started to show significantly lower scores. At the same time, a few old submissions still show old higher scores that makes my submission list look like a mess. Since we already spent weeks on building our models, it is obvious that this situation doesn't affect motivation in a positive way.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2694272,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "03/13/2024 00:34:56",
          "content": "<p>All of the previous of the Notebooks that were scored against the old test data were invalidated. We didn't re-run any automatically with the new test data. Did you re-submit these? I'm curious to understand this.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2694275,
              "author_name": "kononenko",
              "author_url": "",
              "post_date": "03/13/2024 00:46:30",
              "content": "<p>I only resubmitted a few notebooks and found that their scores dropped tremendously, like x2. My guess is that may be the test data have been shuffled and the model somehow expected they're sorted by date or <code>case_id</code>, still trying to figure out the reason.</p>\n<p>All the other submissions are now marked as Error/Succeeded, that I don't know how to interpret:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2099265%2F31ec60d703d57a815069305627cbf90d%2FScreenshot%202024-03-12%20at%205.44.59%20PM.png?generation=1710290723618474&amp;alt=media\"></p>\n<p>If you confirm they were not rescored, then it is easier to simply forget about those LB scores and start from scratch.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2695647,
              "author_name": "octaviograu",
              "author_url": "",
              "post_date": "03/13/2024 19:33:49",
              "content": "<p>I resubmitted a notebook exactly the same and worked <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> <a href=\"https://www.kaggle.com/kononenko\" target=\"_blank\">@kononenko</a> . Other notebooks do not work and need to be rewritten (weekend work for me! :))</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2695022,
      "author_name": "jacobyjaeger",
      "author_url": "",
      "post_date": "03/13/2024 12:17:02",
      "content": "<p>More competitions is generally better, but when a competition turns into a gamble, it affects not only the competition itself but the value of ranks and medals on the platform. Competition hosts benefit from the presence of competitors who are willing to work for little expected monetary reward in exchange for these platform accolades, so I think it is not unreasonable that hosts be held to some standards of security and robustness. </p>\n<p>Would Kaggle not be willing to consider adopting some objective rules for the metrics it allows to avoid the most transparent hackability risks? The problems with this competition could have been avoided if you required that metrics be monotonic with respect to sample scores.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2695064,
          "author_name": "mitsuyasuhoshino",
          "author_url": "",
          "post_date": "03/13/2024 12:59:12",
          "content": "<p>I fully understand both the author's and your points. And one of the common rights given to Kagglers is not to participate in competitions they don't like (since they've restarted, there's no going back).</p>",
          "votes": null,
          "replies": [
            {
              "id": 2695073,
              "author_name": "jacobyjaeger",
              "author_url": "",
              "post_date": "03/13/2024 13:12:44",
              "content": "<p>As I just said, when points and medals are distributed randomly it degrades their value. My post is about Kaggle policy going forward, not just this competition.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2695088,
                  "author_name": "mitsuyasuhoshino",
                  "author_url": "",
                  "post_date": "03/13/2024 13:25:39",
                  "content": "<p>As the former Kaggler author said, I understand that Kaggle competitions cannot eliminate such imperfections. Also, even if a competition is found to have a flaw, it probably won't end mid-way. I will not be participating in this competition, but I will keep an eye on how it goes.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2695094,
                      "author_name": "jacobyjaeger",
                      "author_url": "",
                      "post_date": "03/13/2024 13:29:56",
                      "content": "<p>Hackability that arises from metrics that are not monotonic wrt sample scores can be eliminated by requiring that chosen metrics are monotonic wrt sample scores. I did not suggest ending this competition mid-way.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2695649,
      "author_name": "octaviograu",
      "author_url": "",
      "post_date": "03/13/2024 19:35:05",
      "content": "<p>Thanks for posting <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> . In my case, I am agree with you. For me it is an honor to participate in Kaggle competition and I understand if a competition has issues. There are worse things in life, even I also understand that we put a lot of effort in them. We all need to understand and be grateful for the massive amount of data and problems we can work on here. Cheers!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2695887,
      "author_name": "romandovega",
      "author_url": "",
      "post_date": "03/13/2024 23:26:53",
      "content": "<p>I agree <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> that this competition’s dataset is very unique for Kaggle. It is very realistic in comparison to what data scientists deal with in the “real” tabular world.  Data isn’t always clean and is often messy like this one … as i mentioned <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/482636\" target=\"_blank\">here</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2696273,
      "author_name": "narsil",
      "author_url": "",
      "post_date": "03/14/2024 08:04:13",
      "content": "<blockquote>\n  <p>The one thing that really struck me when I was on the \"other side\" was just how difficult it is to launch a successful competition. There are so many things that can go wrong with the data, so many ways leakage can sneak through, and many many other issues that can arise. </p>\n</blockquote>\n<p>You mention only one side of the equation here. The other one is that you have thousands of super-smart, motivated people who are looking at the competition and will point out EVERY error there is. Which is a good thing. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2697405,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "03/14/2024 21:58:07",
          "content": "<p>I've seen quite a few brilliant researches have flaws in their data discovered by the Kaggle community.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2699850,
              "author_name": "devbilalkhan",
              "author_url": "",
              "post_date": "03/16/2024 07:40:03",
              "content": "<p>I am currently in a competition and have seen participants brought these points about the flaws in the given data. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2699811,
      "author_name": "pourchot",
      "author_url": "",
      "post_date": "03/16/2024 06:45:31",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> , Thank you for your work that is not always simple.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2703270,
      "author_name": "belluri007",
      "author_url": "",
      "post_date": "03/18/2024 04:52:37",
      "content": "<p>Yes, I do agree with <a href=\"https://www.kaggle.com/INVERSION\" target=\"_blank\">@INVERSION</a> that the Home credit has provided a multi-table data which we often don't see not only in Kaggle competitions but also the ones that I have seen outside. This is how the real life problems in AI look like and we spend lot of time on data. </p>\n<p>Here is my approach which I would like to share with this community,</p>\n<ol>\n<li><p>The very first step is to understand the data (all the CVS/Parquet files) holistically. For that one needs to understand the relationship between the data in different files. Look at these files like tables in a RDBMS and establish relationship between them. This is very time consuming process since we need to understand the what each column name means and its relevance in solving this problem.</p></li>\n<li><p>Domain Knowledge - This will help in understanding the data and process flow better. There are good videos and text material on Internet. </p></li>\n<li><p>Its good use a drawing board to draw these files (as tables) and try to visually link them.</p></li>\n<li><p>EDA (of course) is an important task and I don't need to tell you experts. </p></li>\n<li><p>Divide the TRAIN data into Train (70%)  and Validate (30%) and apply on different ML models.</p></li>\n<li><p>Publish and see the scores on leader board. </p></li>\n<li><p>Last but not the least, have a team of 2 to 3 people to brainstorm ideas. </p></li>\n</ol>\n<p>Happy Winning !!! </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2694213": "As I've been reading the question, comments, criticisms, and other feedback about the metric and test dataset update, a few things have come to mind that I'd like to share.\n\nI actively competed in Kaggle competitions for about 5 years before I joined Kaggle 7 years ago. The one thing that really struck me when I was on the \"other side\" was just how difficult it is to launch a successful competition. There are so many things that can go wrong with the data, so many ways leakage can sneak through, and many many other issues that can arise. But even with the (often embarrassing) hiccups that have occurred through the years, I believe the Kaggle platform has had an amazing track record of giving the community fun, interesting, and challenging competitions, year after year.\n\nWe often have to make hard decisions when it comes to launching competitions. There are competitions that come to us in great shape. There are competitions that just aren't feasible that we turn down. And then there are quite a few in the middle, where we have to make \"design decisions\" that might not be optimally tuned for a host or the community. We generally take the stance that it is better to launch an interesting competition with some rough edges than not to launch it at all.\n\nI believe that Home Credit has provided the community with a very interesting and valuable multi-table dataset, something we don't see very often on Kaggle these days. My hope is that, for those who are interested in this type of problem, you can look past the rough edges and still find your participation valuable. Some will choose not to, which is perfectly fine as well. I personally feel the weight when individuals in the community are frustrated or feel like they've wasted their time in a competition. I sincerely hope you'll stick around, and look forward to the interesting approaches for tackling a tabular problem of this complexity.\n\nThanks for listening!",
    "2694266": "I think the main problem here is confusion about what exactly is going on. For example, all of my old submissions now show errors, however, some seems to have been rescored on the new test data, because they started to show significantly lower scores. At the same time, a few old submissions still show old higher scores that makes my submission list look like a mess. Since we already spent weeks on building our models, it is obvious that this situation doesn't affect motivation in a positive way.",
    "2694272": "All of the previous of the Notebooks that were scored against the old test data were invalidated. We didn't re-run any automatically with the new test data. Did you re-submit these? I'm curious to understand this.",
    "2694275": "I only resubmitted a few notebooks and found that their scores dropped tremendously, like x2. My guess is that may be the test data have been shuffled and the model somehow expected they're sorted by date or `case_id`, still trying to figure out the reason.\n\nAll the other submissions are now marked as Error/Succeeded, that I don't know how to interpret:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2099265%2F31ec60d703d57a815069305627cbf90d%2FScreenshot%202024-03-12%20at%205.44.59%20PM.png?generation=1710290723618474&alt=media)\n\nIf you confirm they were not rescored, then it is easier to simply forget about those LB scores and start from scratch.",
    "2695022": "More competitions is generally better, but when a competition turns into a gamble, it affects not only the competition itself but the value of ranks and medals on the platform. Competition hosts benefit from the presence of competitors who are willing to work for little expected monetary reward in exchange for these platform accolades, so I think it is not unreasonable that hosts be held to some standards of security and robustness. \n\nWould Kaggle not be willing to consider adopting some objective rules for the metrics it allows to avoid the most transparent hackability risks? The problems with this competition could have been avoided if you required that metrics be monotonic with respect to sample scores.",
    "2695064": "I fully understand both the author's and your points. And one of the common rights given to Kagglers is not to participate in competitions they don't like (since they've restarted, there's no going back).",
    "2695073": "As I just said, when points and medals are distributed randomly it degrades their value. My post is about Kaggle policy going forward, not just this competition.",
    "2695088": "As the former Kaggler author said, I understand that Kaggle competitions cannot eliminate such imperfections. Also, even if a competition is found to have a flaw, it probably won't end mid-way. I will not be participating in this competition, but I will keep an eye on how it goes.",
    "2695094": "Hackability that arises from metrics that are not monotonic wrt sample scores can be eliminated by requiring that chosen metrics are monotonic wrt sample scores. I did not suggest ending this competition mid-way.",
    "2695647": "I resubmitted a notebook exactly the same and worked @inversion @kononenko . Other notebooks do not work and need to be rewritten (weekend work for me! :))",
    "2695649": "Thanks for posting @inversion . In my case, I am agree with you. For me it is an honor to participate in Kaggle competition and I understand if a competition has issues. There are worse things in life, even I also understand that we put a lot of effort in them. We all need to understand and be grateful for the massive amount of data and problems we can work on here. Cheers!",
    "2695887": "I agree @inversion that this competition’s dataset is very unique for Kaggle. It is very realistic in comparison to what data scientists deal with in the “real” tabular world.  Data isn’t always clean and is often messy like this one … as i mentioned [here](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/482636)",
    "2696273": ">  The one thing that really struck me when I was on the \"other side\" was just how difficult it is to launch a successful competition. There are so many things that can go wrong with the data, so many ways leakage can sneak through, and many many other issues that can arise. \n\nYou mention only one side of the equation here. The other one is that you have thousands of super-smart, motivated people who are looking at the competition and will point out EVERY error there is. Which is a good thing.",
    "2697405": "I've seen quite a few brilliant researches have flaws in their data discovered by the Kaggle community.",
    "2699811": "Hi @inversion , Thank you for your work that is not always simple.",
    "2699850": "I am currently in a competition and have seen participants brought these points about the flaws in the given data.",
    "2703270": "Yes, I do agree with @INVERSION that the Home credit has provided a multi-table data which we often don't see not only in Kaggle competitions but also the ones that I have seen outside. This is how the real life problems in AI look like and we spend lot of time on data. \n\nHere is my approach which I would like to share with this community,\n\n1. The very first step is to understand the data (all the CVS/Parquet files) holistically. For that one needs to understand the relationship between the data in different files. Look at these files like tables in a RDBMS and establish relationship between them. This is very time consuming process since we need to understand the what each column name means and its relevance in solving this problem.\n\n2. Domain Knowledge - This will help in understanding the data and process flow better. There are good videos and text material on Internet. \n\n3. Its good use a drawing board to draw these files (as tables) and try to visually link them.\n\n4. EDA (of course) is an important task and I don't need to tell you experts. \n\n5. Divide the TRAIN data into Train (70%)  and Validate (30%) and apply on different ML models.\n\n6. Publish and see the scores on leader board. \n\n7. Last but not the least, have a team of 2 to 3 people to brainstorm ideas. \n\nHappy Winning !!!"
  },
  "source": "meta"
}