{
  "id": 171580,
  "title": "Getting Started with Competitions",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/171580",
  "author_name": "",
  "post_date": "2020-08-01T15:20:47.804640800Z",
  "votes": 10,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Kaggle Competitions can initially be intimidating for the new data science or machine learning enthusiast. That's the reason there are 'knowledge' based competitions like the Titanic dataset, House prediction to help you gain experience in implementing baseline algorithms for such contests and learn how you can prepare submission files and move on with building more robust and accurate models. The OSIC Pulmonary Fibrosis competition is also a very interesting competition which mainly depends on your skills in CV and inferential statistics skills.</p>\n\n<p>So for the newbie here is my advice on how to proceed with predictive modelling competitions:</p>\n\n<p><strong>Inspection</strong>: Unlike what the novice might assume, building a model isn't the first priority, we always start by inspecting the data. A basic data.head(), data.info(), data.describe() can give you the initial impressions.</p>\n\n<p><strong>Understanding the structure</strong>: This is is crucial, datasets are unclean (in the real world), contain many inconsistencies and are sometimes difficult to perceive. Each dataset is different and pre processing of the data is always unique ( Image data, Time-series data, Numerical, Categorical etc.)</p>\n\n<p><strong>Data Cleaning</strong>: Removing those NULLs, dummies, duplicates are of utmost priority to prevent your data from being skewed. Take account of your outliers through whisker plots/box-plots.</p>\n\n<p><strong>Data visualization</strong>: This is the most important skill and technique that is highly essential for everyone to know. It helps you to easily infer, perceive and understand with the help of simple plots. Start with bar charts, scatter plots and pie charts, move on to histograms, relplots, lmplots, pair-plots.</p>\n\n<p><strong>Modelling</strong>: It is at this step, when you are clear and when you thoroughly understand the data, that you are ready to think about building a model. Brainstorm through the algorithms you have learnt, try to relate this problem to the problems/examples you have seen before.</p>\n\n<p><strong>Optimize</strong>: Once you have the baseline model, it is time to start optimizing it, try to manually tune the hyper parameters, mess around with your test and validation sets. Later you can move on to Hyper-parameter Optimization techniques like RSCV, GSCV or hyper-opt.</p>\n\n<p><strong>Just do it</strong>: A lot of times, people get intimidated by the notebooks by grand-masters, graphs they see and the complex structure of data. Don't get flustered! Start with what you know, when you get stuck you always have great references here at kaggle, Always try yourself first.</p>\n\n<p><em>Last comment: When stuck while coding, Refer documentation of seaborn, matplotlib and any other sklearn/tensorflow models that you may use. No one can by heart the syntaxes of n-number of libraries and functions. The objective is to learn the techniques and concepts in a fun and competitive way. Not to mug up the syntax!</em></p>\n\n<p>I hope i have reduced a bit of your fear and anxiousness in entering such competitions. Remember we are competitors SECOND, students FIRST !!</p>\n\n<p>If my fellow kagglers agree/disagree with anything i have said, Please feel free to comment!</p>",
  "messages": [
    {
      "id": "954269",
      "postDate": "08/01/2020 15:20:47",
      "content": "<p>Kaggle Competitions can initially be intimidating for the new data science or machine learning enthusiast. That's the reason there are 'knowledge' based competitions like the Titanic dataset, House prediction to help you gain experience in implementing baseline algorithms for such contests and learn how you can prepare submission files and move on with building more robust and accurate models. The OSIC Pulmonary Fibrosis competition is also a very interesting competition which mainly depends on your skills in CV and inferential statistics skills.</p>\n\n<p>So for the newbie here is my advice on how to proceed with predictive modelling competitions:</p>\n\n<p><strong>Inspection</strong>: Unlike what the novice might assume, building a model isn't the first priority, we always start by inspecting the data. A basic data.head(), data.info(), data.describe() can give you the initial impressions.</p>\n\n<p><strong>Understanding the structure</strong>: This is is crucial, datasets are unclean (in the real world), contain many inconsistencies and are sometimes difficult to perceive. Each dataset is different and pre processing of the data is always unique ( Image data, Time-series data, Numerical, Categorical etc.)</p>\n\n<p><strong>Data Cleaning</strong>: Removing those NULLs, dummies, duplicates are of utmost priority to prevent your data from being skewed. Take account of your outliers through whisker plots/box-plots.</p>\n\n<p><strong>Data visualization</strong>: This is the most important skill and technique that is highly essential for everyone to know. It helps you to easily infer, perceive and understand with the help of simple plots. Start with bar charts, scatter plots and pie charts, move on to histograms, relplots, lmplots, pair-plots.</p>\n\n<p><strong>Modelling</strong>: It is at this step, when you are clear and when you thoroughly understand the data, that you are ready to think about building a model. Brainstorm through the algorithms you have learnt, try to relate this problem to the problems/examples you have seen before.</p>\n\n<p><strong>Optimize</strong>: Once you have the baseline model, it is time to start optimizing it, try to manually tune the hyper parameters, mess around with your test and validation sets. Later you can move on to Hyper-parameter Optimization techniques like RSCV, GSCV or hyper-opt.</p>\n\n<p><strong>Just do it</strong>: A lot of times, people get intimidated by the notebooks by grand-masters, graphs they see and the complex structure of data. Don't get flustered! Start with what you know, when you get stuck you always have great references here at kaggle, Always try yourself first.</p>\n\n<p><em>Last comment: When stuck while coding, Refer documentation of seaborn, matplotlib and any other sklearn/tensorflow models that you may use. No one can by heart the syntaxes of n-number of libraries and functions. The objective is to learn the techniques and concepts in a fun and competitive way. Not to mug up the syntax!</em></p>\n\n<p>I hope i have reduced a bit of your fear and anxiousness in entering such competitions. Remember we are competitors SECOND, students FIRST !!</p>\n\n<p>If my fellow kagglers agree/disagree with anything i have said, Please feel free to comment!</p>",
      "rawMarkdown": "Kaggle Competitions can initially be intimidating for the new data science or machine learning enthusiast. That's the reason there are 'knowledge' based competitions like the Titanic dataset, House prediction to help you gain experience in implementing baseline algorithms for such contests and learn how you can prepare submission files and move on with building more robust and accurate models. The OSIC Pulmonary Fibrosis competition is also a very interesting competition which mainly depends on your skills in CV and inferential statistics skills.\n\nSo for the newbie here is my advice on how to proceed with predictive modelling competitions:\n\n**Inspection**: Unlike what the novice might assume, building a model isn't the first priority, we always start by inspecting the data. A basic data.head(), data.info(), data.describe() can give you the initial impressions.\n\n**Understanding the structure**: This is is crucial, datasets are unclean (in the real world), contain many inconsistencies and are sometimes difficult to perceive. Each dataset is different and pre processing of the data is always unique ( Image data, Time-series data, Numerical, Categorical etc.)\n\n**Data Cleaning**: Removing those NULLs, dummies, duplicates are of utmost priority to prevent your data from being skewed. Take account of your outliers through whisker plots/box-plots.\n\n**Data visualization**: This is the most important skill and technique that is highly essential for everyone to know. It helps you to easily infer, perceive and understand with the help of simple plots. Start with bar charts, scatter plots and pie charts, move on to histograms, relplots, lmplots, pair-plots.\n\n**Modelling**: It is at this step, when you are clear and when you thoroughly understand the data, that you are ready to think about building a model. Brainstorm through the algorithms you have learnt, try to relate this problem to the problems/examples you have seen before.\n\n**Optimize**: Once you have the baseline model, it is time to start optimizing it, try to manually tune the hyper parameters, mess around with your test and validation sets. Later you can move on to Hyper-parameter Optimization techniques like RSCV, GSCV or hyper-opt.\n\n**Just do it**: A lot of times, people get intimidated by the notebooks by grand-masters, graphs they see and the complex structure of data. Don't get flustered! Start with what you know, when you get stuck you always have great references here at kaggle, Always try yourself first.\n\n*Last comment: When stuck while coding, Refer documentation of seaborn, matplotlib and any other sklearn/tensorflow models that you may use. No one can by heart the syntaxes of n-number of libraries and functions. The objective is to learn the techniques and concepts in a fun and competitive way. Not to mug up the syntax!*\n\nI hope i have reduced a bit of your fear and anxiousness in entering such competitions. Remember we are competitors SECOND, students FIRST !!\n\nIf my fellow kagglers agree/disagree with anything i have said, Please feel free to comment!",
      "votes": null
    },
    {
      "id": "954431",
      "postDate": "08/01/2020 18:11:06",
      "content": "<p>One of the trickiest things is that we all start by forking a notebook written by someone else in their style. The very best are able to write in such a way that you can pick it up and learn but often that's not the case. Look for a 'starter' that is the clearest to you in style and methods not just the best scoring or even the one with the most upvotes (it will become 'your' kernel very quickly). </p>\n\n<p>If you understand each of the processes in your kernel then you can much more easily make (and debug!) the small changes that people propose throughout the competition or interesting ideas that you see implemented in other kernels.</p>",
      "rawMarkdown": "One of the trickiest things is that we all start by forking a notebook written by someone else in their style. The very best are able to write in such a way that you can pick it up and learn but often that's not the case. Look for a 'starter' that is the clearest to you in style and methods not just the best scoring or even the one with the most upvotes (it will become 'your' kernel very quickly). \n\nIf you understand each of the processes in your kernel then you can much more easily make (and debug!) the small changes that people propose throughout the competition or interesting ideas that you see implemented in other kernels.",
      "votes": null
    },
    {
      "id": "954828",
      "postDate": "08/02/2020 06:09:07",
      "content": "<p>That is very true <a href=\"/jameschapman19\">@jameschapman19</a> ! Totally agree.  It soon becomes an involuntary habit to start off by forking someone's notebook! I guess the best a person can do, is to avoid that temptation and try as much as possible to implement his/her own ideas and later compare those with approaches of other kernels. That way at least you'll learn from your mistakes, instead of blindly having code handed on a silver platter. 😄 </p>",
      "rawMarkdown": "That is very true @jameschapman19 ! Totally agree.  It soon becomes an involuntary habit to start off by forking someone's notebook! I guess the best a person can do, is to avoid that temptation and try as much as possible to implement his/her own ideas and later compare those with approaches of other kernels. That way at least you'll learn from your mistakes, instead of blindly having code handed on a silver platter. 😄",
      "votes": null
    },
    {
      "id": "956306",
      "postDate": "08/03/2020 11:56:50",
      "content": "<p>Thank you very much. It boosts my confidence.</p>",
      "rawMarkdown": "Thank you very much. It boosts my confidence.",
      "votes": null
    },
    {
      "id": "994913",
      "postDate": "09/02/2020 04:17:31",
      "content": "<p>Really informative <a href=\"https://www.kaggle.com/darkknight98\" target=\"_blank\">@darkknight98</a>! It really helped me, being a beginner myself! Your work is absolutely great</p>",
      "rawMarkdown": "Really informative @darkknight98! It really helped me, being a beginner myself! Your work is absolutely great",
      "votes": null
    },
    {
      "id": "995039",
      "postDate": "09/02/2020 06:22:06",
      "content": "<p>I totally agree <a href=\"https://www.kaggle.com/darkknight98\" target=\"_blank\">@darkknight98</a>! Very informative and useful 😄</p>",
      "rawMarkdown": "I totally agree @darkknight98! Very informative and useful 😄",
      "votes": null
    },
    {
      "id": "995052",
      "postDate": "09/02/2020 06:29:04",
      "content": "<p>Could you please let me know where i can learn Hyper-Parameter Optimization techniques?</p>",
      "rawMarkdown": "Could you please let me know where i can learn Hyper-Parameter Optimization techniques?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 994913,
      "author_name": "shyam982",
      "author_url": "",
      "post_date": "09/02/2020 04:17:31",
      "content": "<p>Really informative <a href=\"https://www.kaggle.com/darkknight98\" target=\"_blank\">@darkknight98</a>! It really helped me, being a beginner myself! Your work is absolutely great</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 995039,
      "author_name": "specters",
      "author_url": "",
      "post_date": "09/02/2020 06:22:06",
      "content": "<p>I totally agree <a href=\"https://www.kaggle.com/darkknight98\" target=\"_blank\">@darkknight98</a>! Very informative and useful 😄</p>",
      "votes": null,
      "replies": [
        {
          "id": 995052,
          "author_name": "specters",
          "author_url": "",
          "post_date": "09/02/2020 06:29:04",
          "content": "<p>Could you please let me know where i can learn Hyper-Parameter Optimization techniques?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 954431,
      "author_name": "jameschapman19",
      "author_url": "",
      "post_date": "08/01/2020 18:11:06",
      "content": "<p>One of the trickiest things is that we all start by forking a notebook written by someone else in their style. The very best are able to write in such a way that you can pick it up and learn but often that's not the case. Look for a 'starter' that is the clearest to you in style and methods not just the best scoring or even the one with the most upvotes (it will become 'your' kernel very quickly). </p>\n\n<p>If you understand each of the processes in your kernel then you can much more easily make (and debug!) the small changes that people propose throughout the competition or interesting ideas that you see implemented in other kernels.</p>",
      "votes": null,
      "replies": [
        {
          "id": 954828,
          "author_name": "darkknight98",
          "author_url": "",
          "post_date": "08/02/2020 06:09:07",
          "content": "<p>That is very true <a href=\"/jameschapman19\">@jameschapman19</a> ! Totally agree.  It soon becomes an involuntary habit to start off by forking someone's notebook! I guess the best a person can do, is to avoid that temptation and try as much as possible to implement his/her own ideas and later compare those with approaches of other kernels. That way at least you'll learn from your mistakes, instead of blindly having code handed on a silver platter. 😄 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 956306,
      "author_name": "sauravsolanki",
      "author_url": "",
      "post_date": "08/03/2020 11:56:50",
      "content": "<p>Thank you very much. It boosts my confidence.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "954269": "Kaggle Competitions can initially be intimidating for the new data science or machine learning enthusiast. That's the reason there are 'knowledge' based competitions like the Titanic dataset, House prediction to help you gain experience in implementing baseline algorithms for such contests and learn how you can prepare submission files and move on with building more robust and accurate models. The OSIC Pulmonary Fibrosis competition is also a very interesting competition which mainly depends on your skills in CV and inferential statistics skills.\n\nSo for the newbie here is my advice on how to proceed with predictive modelling competitions:\n\n**Inspection**: Unlike what the novice might assume, building a model isn't the first priority, we always start by inspecting the data. A basic data.head(), data.info(), data.describe() can give you the initial impressions.\n\n**Understanding the structure**: This is is crucial, datasets are unclean (in the real world), contain many inconsistencies and are sometimes difficult to perceive. Each dataset is different and pre processing of the data is always unique ( Image data, Time-series data, Numerical, Categorical etc.)\n\n**Data Cleaning**: Removing those NULLs, dummies, duplicates are of utmost priority to prevent your data from being skewed. Take account of your outliers through whisker plots/box-plots.\n\n**Data visualization**: This is the most important skill and technique that is highly essential for everyone to know. It helps you to easily infer, perceive and understand with the help of simple plots. Start with bar charts, scatter plots and pie charts, move on to histograms, relplots, lmplots, pair-plots.\n\n**Modelling**: It is at this step, when you are clear and when you thoroughly understand the data, that you are ready to think about building a model. Brainstorm through the algorithms you have learnt, try to relate this problem to the problems/examples you have seen before.\n\n**Optimize**: Once you have the baseline model, it is time to start optimizing it, try to manually tune the hyper parameters, mess around with your test and validation sets. Later you can move on to Hyper-parameter Optimization techniques like RSCV, GSCV or hyper-opt.\n\n**Just do it**: A lot of times, people get intimidated by the notebooks by grand-masters, graphs they see and the complex structure of data. Don't get flustered! Start with what you know, when you get stuck you always have great references here at kaggle, Always try yourself first.\n\n*Last comment: When stuck while coding, Refer documentation of seaborn, matplotlib and any other sklearn/tensorflow models that you may use. No one can by heart the syntaxes of n-number of libraries and functions. The objective is to learn the techniques and concepts in a fun and competitive way. Not to mug up the syntax!*\n\nI hope i have reduced a bit of your fear and anxiousness in entering such competitions. Remember we are competitors SECOND, students FIRST !!\n\nIf my fellow kagglers agree/disagree with anything i have said, Please feel free to comment!",
    "954431": "One of the trickiest things is that we all start by forking a notebook written by someone else in their style. The very best are able to write in such a way that you can pick it up and learn but often that's not the case. Look for a 'starter' that is the clearest to you in style and methods not just the best scoring or even the one with the most upvotes (it will become 'your' kernel very quickly). \n\nIf you understand each of the processes in your kernel then you can much more easily make (and debug!) the small changes that people propose throughout the competition or interesting ideas that you see implemented in other kernels.",
    "954828": "That is very true @jameschapman19 ! Totally agree.  It soon becomes an involuntary habit to start off by forking someone's notebook! I guess the best a person can do, is to avoid that temptation and try as much as possible to implement his/her own ideas and later compare those with approaches of other kernels. That way at least you'll learn from your mistakes, instead of blindly having code handed on a silver platter. 😄",
    "956306": "Thank you very much. It boosts my confidence.",
    "994913": "Really informative @darkknight98! It really helped me, being a beginner myself! Your work is absolutely great",
    "995039": "I totally agree @darkknight98! Very informative and useful 😄",
    "995052": "Could you please let me know where i can learn Hyper-Parameter Optimization techniques?"
  },
  "source": "meta"
}