{
  "id": 475184,
  "title": "First time participating in Non-Playground competition. Any helpful advice or tips?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/475184",
  "author_name": "Anzar",
  "post_date": "2024-02-07T12:44:08.128000",
  "votes": 28,
  "comment_count": 17,
  "views": 0,
  "content": "<p>I am overwhelmed with the size and complexity of the dataset that I am unable to even start in the first place. How should I approach coded competitions? </p>",
  "messages": [
    {
      "id": 2641330,
      "postDate": "2024-02-07T12:44:08.130Z",
      "content": "<p>I am overwhelmed with the size and complexity of the dataset that I am unable to even start in the first place. How should I approach coded competitions? </p>",
      "rawMarkdown": "I am overwhelmed with the size and complexity of the dataset that I am unable to even start in the first place. How should I approach coded competitions? ",
      "votes": 28
    },
    {
      "id": 2641447,
      "postDate": "2024-02-07T13:43:05.220Z",
      "content": "<p><a href=\"https://www.kaggle.com/anzarwani2\" target=\"_blank\">@anzarwani2</a> some suggestions-</p>\n<ol>\n<li>Focus on learning and less on results </li>\n<li>Work with a better PC/ cloud platform that offers better hardware. Kaggle kernels may be used only to submit models and nothing much more. Please note that Kaggle kernels can be connected to Google Colab now without additional efforts. </li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8273630%2F7c1954f50bd66511c18ed635b3868a86%2FColab.png?generation=1707405513741195&amp;alt=media\"></p>\n<ol>\n<li>Team up well and strategize the team to get the best collective result</li>\n<li>Focus on feature engineering -this is a good strategy to acquire a good score</li>\n<li>Plan your time well and stay active in the challenge for 3 months </li>\n<li>Use ideas from public forums but develop your own pipeline and incorporate these ideas as deemed necessary- this is a good way to acquire control on the work.</li>\n<li>Contribute to the public forums in your best capability</li>\n<li>Learn from other solutions at the end of the competition- this is a valuable learning for future experience. </li>\n</ol>\n<p>All the best!</p>",
      "rawMarkdown": "@anzarwani2 some suggestions-\n1. Focus on learning and less on results \n2. Work with a better PC/ cloud platform that offers better hardware. Kaggle kernels may be used only to submit models and nothing much more. Please note that Kaggle kernels can be connected to Google Colab now without additional efforts. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8273630%2F7c1954f50bd66511c18ed635b3868a86%2FColab.png?generation=1707405513741195&alt=media)\n\n3. Team up well and strategize the team to get the best collective result\n4. Focus on feature engineering -this is a good strategy to acquire a good score\n5. Plan your time well and stay active in the challenge for 3 months \n6. Use ideas from public forums but develop your own pipeline and incorporate these ideas as deemed necessary- this is a good way to acquire control on the work.\n7. Contribute to the public forums in your best capability\n8. Learn from other solutions at the end of the competition- this is a valuable learning for future experience. \n\nAll the best!",
      "votes": 24,
      "replies": [
        {
          "id": 2642460,
          "postDate": "2024-02-08T06:54:20.733Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> for your suggestions and advice!</p>",
          "rawMarkdown": "Thank you @ravi20076 for your suggestions and advice!",
          "votes": 1
        },
        {
          "id": 2643094,
          "postDate": "2024-02-08T16:03:40.883Z",
          "content": "<p>but to have better environment than kaggle notebooks you should probably pay for this, what if i dont have this opportunity?:( What can you recommend instead of kaggle notebooks?</p>",
          "rawMarkdown": "but to have better environment than kaggle notebooks you should probably pay for this, what if i dont have this opportunity?:( What can you recommend instead of kaggle notebooks?",
          "votes": 2,
          "replies": [
            {
              "id": 2645220,
              "postDate": "2024-02-10T04:35:49.810Z",
              "content": "<p><a href=\"https://www.kaggle.com/nazariykarpov\" target=\"_blank\">@nazariykarpov</a> this is a big problem for several participants in this genre. I suggest the below for you-</p>\n<ol>\n<li>Be very choosy regarding your participation. You surely do not want to participate where you will max-out on your resources and not be able to contribute meaningfully</li>\n<li>Form a team with another person with higher resources- in this case, you can contribute meaningfully with your ideas and other member can focus on implementing them. This is usually possible with friends and close colleagues, so team up well</li>\n</ol>",
              "rawMarkdown": "@nazariykarpov this is a big problem for several participants in this genre. I suggest the below for you-\n1. Be very choosy regarding your participation. You surely do not want to participate where you will max-out on your resources and not be able to contribute meaningfully\n2. Form a team with another person with higher resources- in this case, you can contribute meaningfully with your ideas and other member can focus on implementing them. This is usually possible with friends and close colleagues, so team up well",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2641476,
      "postDate": "2024-02-07T14:00:31.253Z",
      "content": "<p>Regarding the dataset, my approach is to spend quite an amount of time going through each csv or parquet file (you just need one format, 2 options are given). After that I make a map of the relations between those files. Which file is doing what and why etc.., refers to bringing them into context. Far after that I start with the usual cleaning process and preprocessing. Keep in mind that you will always end up going back to the data. Even you squeeze out the best model in the beginning. It is based on data, and data is always the bottleneck. Don't run into the mistake and try to model first (because it makes the obvious impact score-wise, but only in the short hand). Only if you have done the job with knowing what you are dealing with and fully understanding the data, THEN other things should come into play. My personal timeline I plan to keep with is: </p>\n<ul>\n<li>First month: EDA, understanding data, roadmap of obstacles</li>\n<li>Second month: Preprocessing with rudimentary models</li>\n<li>Last month: Focus on models and Hyper-parameter-Tuning</li>\n</ul>\n<p>Hope that helps - have fun!</p>",
      "rawMarkdown": "Regarding the dataset, my approach is to spend quite an amount of time going through each csv or parquet file (you just need one format, 2 options are given). After that I make a map of the relations between those files. Which file is doing what and why etc.., refers to bringing them into context. Far after that I start with the usual cleaning process and preprocessing. Keep in mind that you will always end up going back to the data. Even you squeeze out the best model in the beginning. It is based on data, and data is always the bottleneck. Don't run into the mistake and try to model first (because it makes the obvious impact score-wise, but only in the short hand). Only if you have done the job with knowing what you are dealing with and fully understanding the data, THEN other things should come into play. My personal timeline I plan to keep with is: \n\n- First month: EDA, understanding data, roadmap of obstacles\n- Second month: Preprocessing with rudimentary models\n- Last month: Focus on models and Hyper-parameter-Tuning\n\nHope that helps - have fun!",
      "votes": 14,
      "replies": [
        {
          "id": 2642461,
          "postDate": "2024-02-08T06:54:50.640Z",
          "content": "<p>This is helpful, thank you.</p>",
          "rawMarkdown": "This is helpful, thank you.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2643138,
      "postDate": "2024-02-08T16:39:06.530Z",
      "content": "<p><a href=\"https://www.kaggle.com/anzarwani2\" target=\"_blank\">@anzarwani2</a> : Hey Anzar,</p>\n<p>Thank you for raising this topic. As mentioned by other experts, use the chance to participate in this contest as a learning opportunity primarily. The results will come down the road as you progress here in Kaggle in a natural fashion.</p>\n<p>Also, I would like to recommed a few resources / tips to make such a learning fruitful for you.</p>\n<p>First of all, you may want to review the publications and videos below</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/discussions/general/429774#2377548\" target=\"_blank\">Efficiently Mastering Kaggle: Tips for Success</a></li>\n<li><a href=\"https://www.kaggle.com/discussions/general/278111\" target=\"_blank\">Kaggle Competitions Getting Started Guide</a></li>\n<li><a href=\"https://www.kaggle.com/discussions/getting-started/185944\" target=\"_blank\">How to win Kaggle Competitions by Kaggle CEO; Anthony Goldbloom</a></li>\n<li><a href=\"https://www.kaggle.com/discussions/questions-and-answers/416104\" target=\"_blank\">How to win Competitions on Kaggle?</a></li>\n</ul>\n<p>On top of that, you may find it useful to review the winning solutions from the past Kaggle competitions as well as learn from inspirations of the winning GMs/Masters/Experts. This <a href=\"https://www.kaggle.com/competitions/2023-kaggle-ai-report/discussion/421036\" target=\"_blank\">post</a> will help you to see the consistent  approach to locating / reviewing such masterpieces. </p>\n<p>So, for this contest specifically, you can review the approaches/solutions for the past Credit Risk- and Risk-assesment-related feature competitions like <a href=\"https://www.kaggle.com/competitions/home-credit-default-risk\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-default-risk</a> , <a href=\"https://www.kaggle.com/competitions/risky-business\" target=\"_blank\">https://www.kaggle.com/competitions/risky-business</a> and similar (you can find more by searching the competition archives with the query like, <a href=\"https://www.kaggle.com/competitions?sortOption=reward&amp;searchQuery=credit+risk)\" target=\"_blank\">https://www.kaggle.com/competitions?sortOption=reward&amp;searchQuery=credit+risk)</a>.</p>\n<p>You should also take into account this contest to operate with the BigData-scale dataset. Therefore, you need to master the skills related to wrangling and processing BigData in Kaggle notebooks under the Code Competition restrictions (the post <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/475485\" target=\"_blank\">Best Practices of Handling BigData-Scale Datasets in Code Competitions</a> could inspire you to utilize some of the best practices of this sort).</p>\n<p>On the other side, there are some hidden issues on the way. The respective discussion <a href=\"https://www.kaggle.com/discussions/general/395168\" target=\"_blank\">thead</a> may give you a good overview of the pitfalls/drawbacks.</p>\n<p>I hope it is helpful. Happy Kaggling!</p>",
      "rawMarkdown": "@anzarwani2 : Hey Anzar,\n\nThank you for raising this topic. As mentioned by other experts, use the chance to participate in this contest as a learning opportunity primarily. The results will come down the road as you progress here in Kaggle in a natural fashion.\n\nAlso, I would like to recommed a few resources / tips to make such a learning fruitful for you.\n\nFirst of all, you may want to review the publications and videos below\n\n- [Efficiently Mastering Kaggle: Tips for Success](https://www.kaggle.com/discussions/general/429774#2377548)\n- [Kaggle Competitions Getting Started Guide](https://www.kaggle.com/discussions/general/278111)\n- [How to win Kaggle Competitions by Kaggle CEO; Anthony Goldbloom](https://www.kaggle.com/discussions/getting-started/185944)\n- [How to win Competitions on Kaggle?](https://www.kaggle.com/discussions/questions-and-answers/416104)\n\nOn top of that, you may find it useful to review the winning solutions from the past Kaggle competitions as well as learn from inspirations of the winning GMs/Masters/Experts. This [post](https://www.kaggle.com/competitions/2023-kaggle-ai-report/discussion/421036) will help you to see the consistent  approach to locating / reviewing such masterpieces. \n\nSo, for this contest specifically, you can review the approaches/solutions for the past Credit Risk- and Risk-assesment-related feature competitions like https://www.kaggle.com/competitions/home-credit-default-risk , https://www.kaggle.com/competitions/risky-business and similar (you can find more by searching the competition archives with the query like, https://www.kaggle.com/competitions?sortOption=reward&searchQuery=credit+risk).\n\nYou should also take into account this contest to operate with the BigData-scale dataset. Therefore, you need to master the skills related to wrangling and processing BigData in Kaggle notebooks under the Code Competition restrictions (the post [Best Practices of Handling BigData-Scale Datasets in Code Competitions](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/475485) could inspire you to utilize some of the best practices of this sort).\n\nOn the other side, there are some hidden issues on the way. The respective discussion [thead](https://www.kaggle.com/discussions/general/395168) may give you a good overview of the pitfalls/drawbacks.\n\nI hope it is helpful. Happy Kaggling!",
      "votes": 9
    },
    {
      "id": 2641526,
      "postDate": "2024-02-07T14:41:19.983Z",
      "content": "<p>From my experience in the AI world I can give you some tips that worked for me in my daily life:</p>\n<ul>\n<li><p>Try to understand the data first. Since it's a tabular data competition, this will be crucial for your future model behaviour. </p></li>\n<li><p>Don't rush to make submissions to the LB, just take your time, understand the data and try to fit a very basic model. A dummy model is a good choice to use it as a baseline. My recommendation is always to try basic models first as they are easier to understand the predictions and the importance of features. </p></li>\n<li><p>Do not use a very complex model at the beginning. You will invest time and resources without a guarantee of obtaining good results.</p></li>\n<li><p>Do not invest so much time on HP tuning. Just tune the HP when you are sure that your features are the right ones.</p></li>\n<li><p>Validate, validate and validate your model performance. This is a key part in Kaggle competitions. Choose the right validation strategy to avoid overfitting. </p></li>\n</ul>\n<p>This are my general tips. Hope it helps and if you have any observation or comment please do not hesitate on reply this post :) </p>",
      "rawMarkdown": "From my experience in the AI world I can give you some tips that worked for me in my daily life:\n\n* Try to understand the data first. Since it's a tabular data competition, this will be crucial for your future model behaviour. \n\n* Don't rush to make submissions to the LB, just take your time, understand the data and try to fit a very basic model. A dummy model is a good choice to use it as a baseline. My recommendation is always to try basic models first as they are easier to understand the predictions and the importance of features. \n\n* Do not use a very complex model at the beginning. You will invest time and resources without a guarantee of obtaining good results.\n\n* Do not invest so much time on HP tuning. Just tune the HP when you are sure that your features are the right ones.\n\n* Validate, validate and validate your model performance. This is a key part in Kaggle competitions. Choose the right validation strategy to avoid overfitting. \n\nThis are my general tips. Hope it helps and if you have any observation or comment please do not hesitate on reply this post :) ",
      "votes": 8,
      "replies": [
        {
          "id": 2642463,
          "postDate": "2024-02-08T06:55:24.310Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/jbowski\" target=\"_blank\">@jbowski</a> </p>",
          "rawMarkdown": "Thank you @jbowski "
        }
      ]
    },
    {
      "id": 2651184,
      "postDate": "2024-02-14T02:46:23.957Z",
      "content": "<p>Here are my tips </p>\n<ol>\n<li><p>Understand the data and task well , this is the first thing you should do as dataset is more complex compared to playground competition , please understand dataset well because you should not be in a state where few weeks down the line where you have a complex model but arent able to answer simple questions on data . <strong>Also set up a strong cross validation strategy locally so that you are able to trust your approach .</strong></p></li>\n<li><p>Use public kernels to get started , dont try to reinvent wheel but understand the kernels in depth , its all about learning efficiently and as much as possible from others. Also having a strong pipeline allows you to iterate quickly through ideas without any delays and hardcoding which is required to be successfull in Kaggle competition</p></li>\n<li><p>Use discussions for ideas like CV vs LB correlation , feature engineering , feature transformations , dataset size reduction , ensembling , Incorporate those ideas in your pipeline and iterate quickly.</p></li>\n<li><p>Focus more on learning and dont worry about final results as long as you are learning through every submission its worth it</p></li>\n<li><p>In last 2-3 weeks spend time ensembling and selecting best models.</p></li>\n<li><p>Select final submission based on local cv and not on public lb</p></li>\n</ol>\n<p>Happy Kaggling :) , Have fun best of luck 😀</p>",
      "rawMarkdown": "Here are my tips \n1. Understand the data and task well , this is the first thing you should do as dataset is more complex compared to playground competition , please understand dataset well because you should not be in a state where few weeks down the line where you have a complex model but arent able to answer simple questions on data . **Also set up a strong cross validation strategy locally so that you are able to trust your approach .**\n\n2. Use public kernels to get started , dont try to reinvent wheel but understand the kernels in depth , its all about learning efficiently and as much as possible from others. Also having a strong pipeline allows you to iterate quickly through ideas without any delays and hardcoding which is required to be successfull in Kaggle competition\n\n3. Use discussions for ideas like CV vs LB correlation , feature engineering , feature transformations , dataset size reduction , ensembling , Incorporate those ideas in your pipeline and iterate quickly.\n\n4. Focus more on learning and dont worry about final results as long as you are learning through every submission its worth it\n\n5. In last 2-3 weeks spend time ensembling and selecting best models.\n\n6. Select final submission based on local cv and not on public lb\n\nHappy Kaggling :) , Have fun best of luck 😀",
      "votes": 5
    },
    {
      "id": 2643264,
      "postDate": "2024-02-08T18:19:46.783Z",
      "content": "<p><a href=\"https://www.kaggle.com/anzarwani2\" target=\"_blank\">@anzarwani2</a>  Always keep an eye on discussions to get some great ideas and use all holy submissions carefully and consistently.  </p>",
      "rawMarkdown": "@anzarwani2  Always keep an eye on discussions to get some great ideas and use all holy submissions carefully and consistently.  ",
      "votes": 3
    },
    {
      "id": 2641461,
      "postDate": "2024-02-07T13:50:25.637Z",
      "content": "<p>I'm pretty new to competitions, too. I like to monitor the Code section of the competition to see what others have tried so I do not waste time pursuing a mediocre solution. It's helpful to pay attention to the public scores in these notebooks. It is early in the competition, so you will see a rapid improvement in the coming weeks. If you are encountering issues with the size and complexity of data, it is likely others are also experiencing that issue. Focus on learning from what others have done to get a decent score on the leaderboard before you craft a unique approach. Good luck!</p>",
      "rawMarkdown": "I'm pretty new to competitions, too. I like to monitor the Code section of the competition to see what others have tried so I do not waste time pursuing a mediocre solution. It's helpful to pay attention to the public scores in these notebooks. It is early in the competition, so you will see a rapid improvement in the coming weeks. If you are encountering issues with the size and complexity of data, it is likely others are also experiencing that issue. Focus on learning from what others have done to get a decent score on the leaderboard before you craft a unique approach. Good luck!",
      "votes": 3,
      "replies": [
        {
          "id": 2642467,
          "postDate": "2024-02-08T06:56:29.613Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/seanbearden\" target=\"_blank\">@seanbearden</a> <br>\nI will learn from public notebooks and proceed accordingly. </p>",
          "rawMarkdown": "Thank you @seanbearden \nI will learn from public notebooks and proceed accordingly. "
        },
        {
          "id": 2689465,
          "postDate": "2024-03-09T21:06:28.667Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2641486,
      "postDate": "2024-02-07T14:08:56.123Z",
      "content": "<p>Understand the data. Do some research, preprocessing. And if it is possible team up with someone. For example with me)</p>",
      "rawMarkdown": "Understand the data. Do some research, preprocessing. And if it is possible team up with someone. For example with me)",
      "votes": 2,
      "replies": [
        {
          "id": 2642469,
          "postDate": "2024-02-08T06:57:19.913Z",
          "content": "<p>I will keep that in mind 😏 <a href=\"https://www.kaggle.com/mrsimple07\" target=\"_blank\">@mrsimple07</a> </p>",
          "rawMarkdown": "I will keep that in mind 😏 @mrsimple07 "
        },
        {
          "id": 2642699,
          "postDate": "2024-02-08T10:56:41.873Z",
          "content": "<p>can I team up with you?</p>",
          "rawMarkdown": "can I team up with you?",
          "votes": -1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2641447,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2024-02-07T13:43:05.220000",
      "content": "<p><a href=\"https://www.kaggle.com/anzarwani2\" target=\"_blank\">@anzarwani2</a> some suggestions-</p>\n<ol>\n<li>Focus on learning and less on results </li>\n<li>Work with a better PC/ cloud platform that offers better hardware. Kaggle kernels may be used only to submit models and nothing much more. Please note that Kaggle kernels can be connected to Google Colab now without additional efforts. </li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8273630%2F7c1954f50bd66511c18ed635b3868a86%2FColab.png?generation=1707405513741195&amp;alt=media\"></p>\n<ol>\n<li>Team up well and strategize the team to get the best collective result</li>\n<li>Focus on feature engineering -this is a good strategy to acquire a good score</li>\n<li>Plan your time well and stay active in the challenge for 3 months </li>\n<li>Use ideas from public forums but develop your own pipeline and incorporate these ideas as deemed necessary- this is a good way to acquire control on the work.</li>\n<li>Contribute to the public forums in your best capability</li>\n<li>Learn from other solutions at the end of the competition- this is a valuable learning for future experience. </li>\n</ol>\n<p>All the best!</p>",
      "votes": 24,
      "replies": [
        {
          "id": 2642460,
          "author_name": "Anzar",
          "author_url": "",
          "post_date": "2024-02-08T06:54:20.733000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> for your suggestions and advice!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2643094,
          "author_name": "Nazariy Karpov",
          "author_url": "",
          "post_date": "2024-02-08T16:03:40.883000",
          "content": "<p>but to have better environment than kaggle notebooks you should probably pay for this, what if i dont have this opportunity?:( What can you recommend instead of kaggle notebooks?</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2645220,
              "author_name": "Ravi Ramakrishnan",
              "author_url": "",
              "post_date": "2024-02-10T04:35:49.810000",
              "content": "<p><a href=\"https://www.kaggle.com/nazariykarpov\" target=\"_blank\">@nazariykarpov</a> this is a big problem for several participants in this genre. I suggest the below for you-</p>\n<ol>\n<li>Be very choosy regarding your participation. You surely do not want to participate where you will max-out on your resources and not be able to contribute meaningfully</li>\n<li>Form a team with another person with higher resources- in this case, you can contribute meaningfully with your ideas and other member can focus on implementing them. This is usually possible with friends and close colleagues, so team up well</li>\n</ol>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2641476,
      "author_name": "Etienne Kaiser (郑翊天）",
      "author_url": "",
      "post_date": "2024-02-07T14:00:31.253000",
      "content": "<p>Regarding the dataset, my approach is to spend quite an amount of time going through each csv or parquet file (you just need one format, 2 options are given). After that I make a map of the relations between those files. Which file is doing what and why etc.., refers to bringing them into context. Far after that I start with the usual cleaning process and preprocessing. Keep in mind that you will always end up going back to the data. Even you squeeze out the best model in the beginning. It is based on data, and data is always the bottleneck. Don't run into the mistake and try to model first (because it makes the obvious impact score-wise, but only in the short hand). Only if you have done the job with knowing what you are dealing with and fully understanding the data, THEN other things should come into play. My personal timeline I plan to keep with is: </p>\n<ul>\n<li>First month: EDA, understanding data, roadmap of obstacles</li>\n<li>Second month: Preprocessing with rudimentary models</li>\n<li>Last month: Focus on models and Hyper-parameter-Tuning</li>\n</ul>\n<p>Hope that helps - have fun!</p>",
      "votes": 14,
      "replies": [
        {
          "id": 2642461,
          "author_name": "Anzar",
          "author_url": "",
          "post_date": "2024-02-08T06:54:50.640000",
          "content": "<p>This is helpful, thank you.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2643138,
      "author_name": "Georgii Vyshnia",
      "author_url": "",
      "post_date": "2024-02-08T16:39:06.530000",
      "content": "<p><a href=\"https://www.kaggle.com/anzarwani2\" target=\"_blank\">@anzarwani2</a> : Hey Anzar,</p>\n<p>Thank you for raising this topic. As mentioned by other experts, use the chance to participate in this contest as a learning opportunity primarily. The results will come down the road as you progress here in Kaggle in a natural fashion.</p>\n<p>Also, I would like to recommed a few resources / tips to make such a learning fruitful for you.</p>\n<p>First of all, you may want to review the publications and videos below</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/discussions/general/429774#2377548\" target=\"_blank\">Efficiently Mastering Kaggle: Tips for Success</a></li>\n<li><a href=\"https://www.kaggle.com/discussions/general/278111\" target=\"_blank\">Kaggle Competitions Getting Started Guide</a></li>\n<li><a href=\"https://www.kaggle.com/discussions/getting-started/185944\" target=\"_blank\">How to win Kaggle Competitions by Kaggle CEO; Anthony Goldbloom</a></li>\n<li><a href=\"https://www.kaggle.com/discussions/questions-and-answers/416104\" target=\"_blank\">How to win Competitions on Kaggle?</a></li>\n</ul>\n<p>On top of that, you may find it useful to review the winning solutions from the past Kaggle competitions as well as learn from inspirations of the winning GMs/Masters/Experts. This <a href=\"https://www.kaggle.com/competitions/2023-kaggle-ai-report/discussion/421036\" target=\"_blank\">post</a> will help you to see the consistent  approach to locating / reviewing such masterpieces. </p>\n<p>So, for this contest specifically, you can review the approaches/solutions for the past Credit Risk- and Risk-assesment-related feature competitions like <a href=\"https://www.kaggle.com/competitions/home-credit-default-risk\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-default-risk</a> , <a href=\"https://www.kaggle.com/competitions/risky-business\" target=\"_blank\">https://www.kaggle.com/competitions/risky-business</a> and similar (you can find more by searching the competition archives with the query like, <a href=\"https://www.kaggle.com/competitions?sortOption=reward&amp;searchQuery=credit+risk)\" target=\"_blank\">https://www.kaggle.com/competitions?sortOption=reward&amp;searchQuery=credit+risk)</a>.</p>\n<p>You should also take into account this contest to operate with the BigData-scale dataset. Therefore, you need to master the skills related to wrangling and processing BigData in Kaggle notebooks under the Code Competition restrictions (the post <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/475485\" target=\"_blank\">Best Practices of Handling BigData-Scale Datasets in Code Competitions</a> could inspire you to utilize some of the best practices of this sort).</p>\n<p>On the other side, there are some hidden issues on the way. The respective discussion <a href=\"https://www.kaggle.com/discussions/general/395168\" target=\"_blank\">thead</a> may give you a good overview of the pitfalls/drawbacks.</p>\n<p>I hope it is helpful. Happy Kaggling!</p>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 2641526,
      "author_name": "JBowski",
      "author_url": "",
      "post_date": "2024-02-07T14:41:19.983000",
      "content": "<p>From my experience in the AI world I can give you some tips that worked for me in my daily life:</p>\n<ul>\n<li><p>Try to understand the data first. Since it's a tabular data competition, this will be crucial for your future model behaviour. </p></li>\n<li><p>Don't rush to make submissions to the LB, just take your time, understand the data and try to fit a very basic model. A dummy model is a good choice to use it as a baseline. My recommendation is always to try basic models first as they are easier to understand the predictions and the importance of features. </p></li>\n<li><p>Do not use a very complex model at the beginning. You will invest time and resources without a guarantee of obtaining good results.</p></li>\n<li><p>Do not invest so much time on HP tuning. Just tune the HP when you are sure that your features are the right ones.</p></li>\n<li><p>Validate, validate and validate your model performance. This is a key part in Kaggle competitions. Choose the right validation strategy to avoid overfitting. </p></li>\n</ul>\n<p>This are my general tips. Hope it helps and if you have any observation or comment please do not hesitate on reply this post :) </p>",
      "votes": 8,
      "replies": [
        {
          "id": 2642463,
          "author_name": "Anzar",
          "author_url": "",
          "post_date": "2024-02-08T06:55:24.310000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/jbowski\" target=\"_blank\">@jbowski</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2651184,
      "author_name": "Athar Sayed",
      "author_url": "",
      "post_date": "2024-02-14T02:46:23.957000",
      "content": "<p>Here are my tips </p>\n<ol>\n<li><p>Understand the data and task well , this is the first thing you should do as dataset is more complex compared to playground competition , please understand dataset well because you should not be in a state where few weeks down the line where you have a complex model but arent able to answer simple questions on data . <strong>Also set up a strong cross validation strategy locally so that you are able to trust your approach .</strong></p></li>\n<li><p>Use public kernels to get started , dont try to reinvent wheel but understand the kernels in depth , its all about learning efficiently and as much as possible from others. Also having a strong pipeline allows you to iterate quickly through ideas without any delays and hardcoding which is required to be successfull in Kaggle competition</p></li>\n<li><p>Use discussions for ideas like CV vs LB correlation , feature engineering , feature transformations , dataset size reduction , ensembling , Incorporate those ideas in your pipeline and iterate quickly.</p></li>\n<li><p>Focus more on learning and dont worry about final results as long as you are learning through every submission its worth it</p></li>\n<li><p>In last 2-3 weeks spend time ensembling and selecting best models.</p></li>\n<li><p>Select final submission based on local cv and not on public lb</p></li>\n</ol>\n<p>Happy Kaggling :) , Have fun best of luck 😀</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 2643264,
      "author_name": "Kishan Vavdara",
      "author_url": "",
      "post_date": "2024-02-08T18:19:46.783000",
      "content": "<p><a href=\"https://www.kaggle.com/anzarwani2\" target=\"_blank\">@anzarwani2</a>  Always keep an eye on discussions to get some great ideas and use all holy submissions carefully and consistently.  </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2641461,
      "author_name": "Sean R.B. Bearden, Ph.D.",
      "author_url": "",
      "post_date": "2024-02-07T13:50:25.637000",
      "content": "<p>I'm pretty new to competitions, too. I like to monitor the Code section of the competition to see what others have tried so I do not waste time pursuing a mediocre solution. It's helpful to pay attention to the public scores in these notebooks. It is early in the competition, so you will see a rapid improvement in the coming weeks. If you are encountering issues with the size and complexity of data, it is likely others are also experiencing that issue. Focus on learning from what others have done to get a decent score on the leaderboard before you craft a unique approach. Good luck!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2642467,
          "author_name": "Anzar",
          "author_url": "",
          "post_date": "2024-02-08T06:56:29.613000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/seanbearden\" target=\"_blank\">@seanbearden</a> <br>\nI will learn from public notebooks and proceed accordingly. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2689465,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-03-09T21:06:28.667000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2641486,
      "author_name": "MrSimple",
      "author_url": "",
      "post_date": "2024-02-07T14:08:56.123000",
      "content": "<p>Understand the data. Do some research, preprocessing. And if it is possible team up with someone. For example with me)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2642469,
          "author_name": "Anzar",
          "author_url": "",
          "post_date": "2024-02-08T06:57:19.913000",
          "content": "<p>I will keep that in mind 😏 <a href=\"https://www.kaggle.com/mrsimple07\" target=\"_blank\">@mrsimple07</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2642699,
          "author_name": "Hajarkhagd",
          "author_url": "",
          "post_date": "2024-02-08T10:56:41.873000",
          "content": "<p>can I team up with you?</p>",
          "votes": -1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2641330": "I am overwhelmed with the size and complexity of the dataset that I am unable to even start in the first place. How should I approach coded competitions? ",
    "2641447": "@anzarwani2 some suggestions-\n1. Focus on learning and less on results \n2. Work with a better PC/ cloud platform that offers better hardware. Kaggle kernels may be used only to submit models and nothing much more. Please note that Kaggle kernels can be connected to Google Colab now without additional efforts. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8273630%2F7c1954f50bd66511c18ed635b3868a86%2FColab.png?generation=1707405513741195&alt=media)\n\n3. Team up well and strategize the team to get the best collective result\n4. Focus on feature engineering -this is a good strategy to acquire a good score\n5. Plan your time well and stay active in the challenge for 3 months \n6. Use ideas from public forums but develop your own pipeline and incorporate these ideas as deemed necessary- this is a good way to acquire control on the work.\n7. Contribute to the public forums in your best capability\n8. Learn from other solutions at the end of the competition- this is a valuable learning for future experience. \n\nAll the best!",
    "2641476": "Regarding the dataset, my approach is to spend quite an amount of time going through each csv or parquet file (you just need one format, 2 options are given). After that I make a map of the relations between those files. Which file is doing what and why etc.., refers to bringing them into context. Far after that I start with the usual cleaning process and preprocessing. Keep in mind that you will always end up going back to the data. Even you squeeze out the best model in the beginning. It is based on data, and data is always the bottleneck. Don't run into the mistake and try to model first (because it makes the obvious impact score-wise, but only in the short hand). Only if you have done the job with knowing what you are dealing with and fully understanding the data, THEN other things should come into play. My personal timeline I plan to keep with is: \n\n- First month: EDA, understanding data, roadmap of obstacles\n- Second month: Preprocessing with rudimentary models\n- Last month: Focus on models and Hyper-parameter-Tuning\n\nHope that helps - have fun!",
    "2643138": "@anzarwani2 : Hey Anzar,\n\nThank you for raising this topic. As mentioned by other experts, use the chance to participate in this contest as a learning opportunity primarily. The results will come down the road as you progress here in Kaggle in a natural fashion.\n\nAlso, I would like to recommed a few resources / tips to make such a learning fruitful for you.\n\nFirst of all, you may want to review the publications and videos below\n\n- [Efficiently Mastering Kaggle: Tips for Success](https://www.kaggle.com/discussions/general/429774#2377548)\n- [Kaggle Competitions Getting Started Guide](https://www.kaggle.com/discussions/general/278111)\n- [How to win Kaggle Competitions by Kaggle CEO; Anthony Goldbloom](https://www.kaggle.com/discussions/getting-started/185944)\n- [How to win Competitions on Kaggle?](https://www.kaggle.com/discussions/questions-and-answers/416104)\n\nOn top of that, you may find it useful to review the winning solutions from the past Kaggle competitions as well as learn from inspirations of the winning GMs/Masters/Experts. This [post](https://www.kaggle.com/competitions/2023-kaggle-ai-report/discussion/421036) will help you to see the consistent  approach to locating / reviewing such masterpieces. \n\nSo, for this contest specifically, you can review the approaches/solutions for the past Credit Risk- and Risk-assesment-related feature competitions like https://www.kaggle.com/competitions/home-credit-default-risk , https://www.kaggle.com/competitions/risky-business and similar (you can find more by searching the competition archives with the query like, https://www.kaggle.com/competitions?sortOption=reward&searchQuery=credit+risk).\n\nYou should also take into account this contest to operate with the BigData-scale dataset. Therefore, you need to master the skills related to wrangling and processing BigData in Kaggle notebooks under the Code Competition restrictions (the post [Best Practices of Handling BigData-Scale Datasets in Code Competitions](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/475485) could inspire you to utilize some of the best practices of this sort).\n\nOn the other side, there are some hidden issues on the way. The respective discussion [thead](https://www.kaggle.com/discussions/general/395168) may give you a good overview of the pitfalls/drawbacks.\n\nI hope it is helpful. Happy Kaggling!",
    "2641526": "From my experience in the AI world I can give you some tips that worked for me in my daily life:\n\n* Try to understand the data first. Since it's a tabular data competition, this will be crucial for your future model behaviour. \n\n* Don't rush to make submissions to the LB, just take your time, understand the data and try to fit a very basic model. A dummy model is a good choice to use it as a baseline. My recommendation is always to try basic models first as they are easier to understand the predictions and the importance of features. \n\n* Do not use a very complex model at the beginning. You will invest time and resources without a guarantee of obtaining good results.\n\n* Do not invest so much time on HP tuning. Just tune the HP when you are sure that your features are the right ones.\n\n* Validate, validate and validate your model performance. This is a key part in Kaggle competitions. Choose the right validation strategy to avoid overfitting. \n\nThis are my general tips. Hope it helps and if you have any observation or comment please do not hesitate on reply this post :) ",
    "2651184": "Here are my tips \n1. Understand the data and task well , this is the first thing you should do as dataset is more complex compared to playground competition , please understand dataset well because you should not be in a state where few weeks down the line where you have a complex model but arent able to answer simple questions on data . **Also set up a strong cross validation strategy locally so that you are able to trust your approach .**\n\n2. Use public kernels to get started , dont try to reinvent wheel but understand the kernels in depth , its all about learning efficiently and as much as possible from others. Also having a strong pipeline allows you to iterate quickly through ideas without any delays and hardcoding which is required to be successfull in Kaggle competition\n\n3. Use discussions for ideas like CV vs LB correlation , feature engineering , feature transformations , dataset size reduction , ensembling , Incorporate those ideas in your pipeline and iterate quickly.\n\n4. Focus more on learning and dont worry about final results as long as you are learning through every submission its worth it\n\n5. In last 2-3 weeks spend time ensembling and selecting best models.\n\n6. Select final submission based on local cv and not on public lb\n\nHappy Kaggling :) , Have fun best of luck 😀",
    "2643264": "@anzarwani2  Always keep an eye on discussions to get some great ideas and use all holy submissions carefully and consistently.  ",
    "2641461": "I'm pretty new to competitions, too. I like to monitor the Code section of the competition to see what others have tried so I do not waste time pursuing a mediocre solution. It's helpful to pay attention to the public scores in these notebooks. It is early in the competition, so you will see a rapid improvement in the coming weeks. If you are encountering issues with the size and complexity of data, it is likely others are also experiencing that issue. Focus on learning from what others have done to get a decent score on the leaderboard before you craft a unique approach. Good luck!",
    "2641486": "Understand the data. Do some research, preprocessing. And if it is possible team up with someone. For example with me)"
  }
}