{
  "id": 486974,
  "title": "How to Navigate Through the Overwhelming Amount of Data in the Consumer Finance Competition and Where to Begin?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/486974",
  "author_name": "",
  "post_date": "2024-03-27T03:48:00.356319800Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hey everyone, I'm feeling a bit overwhelmed by the sheer volume of data in this competition aimed at predicting default risk for consumer finance providers. Can anyone offer guidance on how to effectively tackle such a large dataset and where to begin in developing predictive models to improve financial inclusion for individuals with limited credit history?</p>",
  "messages": [
    {
      "id": "2718261",
      "postDate": "03/27/2024 03:48:00",
      "content": "<p>Hey everyone, I'm feeling a bit overwhelmed by the sheer volume of data in this competition aimed at predicting default risk for consumer finance providers. Can anyone offer guidance on how to effectively tackle such a large dataset and where to begin in developing predictive models to improve financial inclusion for individuals with limited credit history?</p>",
      "rawMarkdown": "Hey everyone, I'm feeling a bit overwhelmed by the sheer volume of data in this competition aimed at predicting default risk for consumer finance providers. Can anyone offer guidance on how to effectively tackle such a large dataset and where to begin in developing predictive models to improve financial inclusion for individuals with limited credit history?",
      "votes": null
    },
    {
      "id": "2719128",
      "postDate": "03/27/2024 14:19:40",
      "content": "<p>Refer to the starter notebook in the code section, it is pinned right on top. I will be a good starting point. Many points are explained in it. Also the use of Polars instead of pandas is helpful, If you are more comfortable with pandas than u can use it to see and check the data for analysis, but for final submission, it will be better to replace pandas with polars for reading the data and data manipulation. Also you can get lot of ideas from various notebooks in public and how they have made their approach. It might take some time to get on with the flow, but once u get the process, everything will go smoothly.</p>",
      "rawMarkdown": "Refer to the starter notebook in the code section, it is pinned right on top. I will be a good starting point. Many points are explained in it. Also the use of Polars instead of pandas is helpful, If you are more comfortable with pandas than u can use it to see and check the data for analysis, but for final submission, it will be better to replace pandas with polars for reading the data and data manipulation. Also you can get lot of ideas from various notebooks in public and how they have made their approach. It might take some time to get on with the flow, but once u get the process, everything will go smoothly.",
      "votes": null
    },
    {
      "id": "2724807",
      "postDate": "03/31/2024 05:33:27",
      "content": "<p>I lot of public code uses polars and handles this extremely well. Even I was of this concern a long while ago, but the quality of  public kernels in this challenge is supreme <a href=\"https://www.kaggle.com/aaditshukla\" target=\"_blank\">@aaditshukla</a>!</p>",
      "rawMarkdown": "I lot of public code uses polars and handles this extremely well. Even I was of this concern a long while ago, but the quality of  public kernels in this challenge is supreme @aaditshukla!",
      "votes": null
    },
    {
      "id": "2728877",
      "postDate": "04/02/2024 12:56:45",
      "content": "<p>Hi, hope this notebook <a href=\"https://www.kaggle.com/code/sani84/start-code-with-pandas-lgb\" target=\"_blank\"></a> will be helpful.</p>",
      "rawMarkdown": "Hi, hope this notebook [](https://www.kaggle.com/code/sani84/start-code-with-pandas-lgb) will be helpful.",
      "votes": null
    },
    {
      "id": "2728878",
      "postDate": "04/02/2024 12:57:35",
      "content": "<p><a href=\"https://www.kaggle.com/code/sani84/start-code-with-pandas-lgb\" target=\"_blank\">https://www.kaggle.com/code/sani84/start-code-with-pandas-lgb</a></p>",
      "rawMarkdown": "https://www.kaggle.com/code/sani84/start-code-with-pandas-lgb",
      "votes": null
    },
    {
      "id": "2739660",
      "postDate": "04/07/2024 07:18:04",
      "content": "<p>thank you for sharing <a href=\"https://www.kaggle.com/sani84\" target=\"_blank\">@sani84</a> </p>",
      "rawMarkdown": "thank you for sharing @sani84",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2719128,
      "author_name": "shreyas9181",
      "author_url": "",
      "post_date": "03/27/2024 14:19:40",
      "content": "<p>Refer to the starter notebook in the code section, it is pinned right on top. I will be a good starting point. Many points are explained in it. Also the use of Polars instead of pandas is helpful, If you are more comfortable with pandas than u can use it to see and check the data for analysis, but for final submission, it will be better to replace pandas with polars for reading the data and data manipulation. Also you can get lot of ideas from various notebooks in public and how they have made their approach. It might take some time to get on with the flow, but once u get the process, everything will go smoothly.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2724807,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "03/31/2024 05:33:27",
      "content": "<p>I lot of public code uses polars and handles this extremely well. Even I was of this concern a long while ago, but the quality of  public kernels in this challenge is supreme <a href=\"https://www.kaggle.com/aaditshukla\" target=\"_blank\">@aaditshukla</a>!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2728877,
      "author_name": "sani84",
      "author_url": "",
      "post_date": "04/02/2024 12:56:45",
      "content": "<p>Hi, hope this notebook <a href=\"https://www.kaggle.com/code/sani84/start-code-with-pandas-lgb\" target=\"_blank\"></a> will be helpful.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2728878,
          "author_name": "sani84",
          "author_url": "",
          "post_date": "04/02/2024 12:57:35",
          "content": "<p><a href=\"https://www.kaggle.com/code/sani84/start-code-with-pandas-lgb\" target=\"_blank\">https://www.kaggle.com/code/sani84/start-code-with-pandas-lgb</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 2739660,
              "author_name": "aaditshukla",
              "author_url": "",
              "post_date": "04/07/2024 07:18:04",
              "content": "<p>thank you for sharing <a href=\"https://www.kaggle.com/sani84\" target=\"_blank\">@sani84</a> </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2718261": "Hey everyone, I'm feeling a bit overwhelmed by the sheer volume of data in this competition aimed at predicting default risk for consumer finance providers. Can anyone offer guidance on how to effectively tackle such a large dataset and where to begin in developing predictive models to improve financial inclusion for individuals with limited credit history?",
    "2719128": "Refer to the starter notebook in the code section, it is pinned right on top. I will be a good starting point. Many points are explained in it. Also the use of Polars instead of pandas is helpful, If you are more comfortable with pandas than u can use it to see and check the data for analysis, but for final submission, it will be better to replace pandas with polars for reading the data and data manipulation. Also you can get lot of ideas from various notebooks in public and how they have made their approach. It might take some time to get on with the flow, but once u get the process, everything will go smoothly.",
    "2724807": "I lot of public code uses polars and handles this extremely well. Even I was of this concern a long while ago, but the quality of  public kernels in this challenge is supreme @aaditshukla!",
    "2728877": "Hi, hope this notebook [](https://www.kaggle.com/code/sani84/start-code-with-pandas-lgb) will be helpful.",
    "2728878": "https://www.kaggle.com/code/sani84/start-code-with-pandas-lgb",
    "2739660": "thank you for sharing @sani84"
  },
  "source": "meta"
}