{
  "id": 53546,
  "title": "Newbie who wants to learn by doing but where to start?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53546",
  "author_name": "",
  "post_date": "2018-04-01T10:17:04.992882600Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi,\nI feel a complete incompetent newbie... :-S\nStruggled my way through the first 3 courses of the Python for Everyone Specialization on Coursera, \nI saw someone mention Kaggle as a good place to practice with simple problems and cases. \nI feel quite overwhelmed by the Kernils and Tutorials voted high in the discussion board. I guess I need more foundation knowledge.... \nI am happy to learn more by watching videos or reading books, but is there a better way to really learn as an ABSOLUTE NEWBIE by DOING?</p>\n\n<p>Any encouragement and recommendations for someone who learns by doing, please?</p>",
  "messages": [
    {
      "id": "307301",
      "postDate": "04/01/2018 10:17:04",
      "content": "<p>Hi,\nI feel a complete incompetent newbie... :-S\nStruggled my way through the first 3 courses of the Python for Everyone Specialization on Coursera, \nI saw someone mention Kaggle as a good place to practice with simple problems and cases. \nI feel quite overwhelmed by the Kernils and Tutorials voted high in the discussion board. I guess I need more foundation knowledge.... \nI am happy to learn more by watching videos or reading books, but is there a better way to really learn as an ABSOLUTE NEWBIE by DOING?</p>\n\n<p>Any encouragement and recommendations for someone who learns by doing, please?</p>",
      "rawMarkdown": "Hi,\nI feel a complete incompetent newbie... :-S\nStruggled my way through the first 3 courses of the Python for Everyone Specialization on Coursera, \nI saw someone mention Kaggle as a good place to practice with simple problems and cases. \nI feel quite overwhelmed by the Kernils and Tutorials voted high in the discussion board. I guess I need more foundation knowledge.... \nI am happy to learn more by watching videos or reading books, but is there a better way to really learn as an ABSOLUTE NEWBIE by DOING?\n\nAny encouragement and recommendations for someone who learns by doing, please?",
      "votes": null
    },
    {
      "id": "307323",
      "postDate": "04/01/2018 11:13:56",
      "content": "<p>Congratulations, you are in the right place to learn by doing :). My recommendation is to start with a smaller dataset (the ones in the Getting Started category). That will allow you to make many experiments fast. Then give yourself the goal of creating from scratch a simple model for the problem (linear or logistic regression for example) and make that work. Then go and try some other algorithms like random forest and gradient boosting. After being able to solve these small problems everything else will be quick. Enjoy.</p>",
      "rawMarkdown": "Congratulations, you are in the right place to learn by doing :). My recommendation is to start with a smaller dataset (the ones in the Getting Started category). That will allow you to make many experiments fast. Then give yourself the goal of creating from scratch a simple model for the problem (linear or logistic regression for example) and make that work. Then go and try some other algorithms like random forest and gradient boosting. After being able to solve these small problems everything else will be quick. Enjoy.",
      "votes": null
    },
    {
      "id": "307425",
      "postDate": "04/01/2018 15:35:40",
      "content": "<p>I agree that learning by doing is the best (at least for me), but nevertheless, I will combine this with studing existing kernels from other people out there. For example, try your best for 4-5 hours on one of the playgrounds or in a past competition and then look at 5-10 of the \"low-level\" kernels (search for keywords like EDA, or tutorial, or baseline). Study what you could do better (or more), maybe try to do it again or just go to the other competition. Repeat.</p>",
      "rawMarkdown": "I agree that learning by doing is the best (at least for me), but nevertheless, I will combine this with studing existing kernels from other people out there. For example, try your best for 4-5 hours on one of the playgrounds or in a past competition and then look at 5-10 of the \"low-level\" kernels (search for keywords like EDA, or tutorial, or baseline). Study what you could do better (or more), maybe try to do it again or just go to the other competition. Repeat.",
      "votes": null
    },
    {
      "id": "308827",
      "postDate": "04/04/2018 07:06:53",
      "content": "<p>Personally i learn the most by writing my own code, as it forces you to think the problem through, otherwise it will not work out. Hard in the beginning but gets easier and easier, also very satisfying, when the program smoothly runs through. </p>\n\n<p>For a start (and basically all the way) begin with building a working pipeline, that is a program that loads in the data, calculates the features (start simple here), runs a predictive model, predicts the test data and writes a submission file that you can submit here (and get a realistic score &gt;0.93). Most kernels have such a pipeline, that you can use for inspiration, however i would recommend to write your own from scratch as this way you learn the most and start to think like a data scientist.</p>\n\n<p>I agree with the others though that this dataset is not great for starting, as the size requires additional considerations, if you dont have a lot of computational ressources. </p>",
      "rawMarkdown": "Personally i learn the most by writing my own code, as it forces you to think the problem through, otherwise it will not work out. Hard in the beginning but gets easier and easier, also very satisfying, when the program smoothly runs through. \n\n For a start (and basically all the way) begin with building a working pipeline, that is a program that loads in the data, calculates the features (start simple here), runs a predictive model, predicts the test data and writes a submission file that you can submit here (and get a realistic score &gt;0.93). Most kernels have such a pipeline, that you can use for inspiration, however i would recommend to write your own from scratch as this way you learn the most and start to think like a data scientist.\n\n I agree with the others though that this dataset is not great for starting, as the size requires additional considerations, if you dont have a lot of computational ressources.",
      "votes": null
    },
    {
      "id": "309072",
      "postDate": "04/04/2018 15:35:09",
      "content": "<p>@ JonW </p>\n\n<p>First up good decision. So Congrats. \nI was in your position about 4 months ago, so I know how you feel. </p>\n\n<p>2 Approaches </p>\n\n<p>1 - Jump headlong into a LIVE competition depending on your area of interest. I started with the Mercari Price Suggestion ( An NLP problem )challenge and I did not even know how to vectorize a sentence when I started this competition. However, going after a single problem gave me the focus to go out and try all kinds of approaches. The learning was incredible. It was frustrating, sleep depriving and down right discouraging as conversations flew thick and fast about things I simply could not follow. I hung in there, kept learning, crashed out of the competition without a rank but the learning was MASSIVE. I now have the confidence to approach and work on NLP problems. I am nowhere near winning or making a dent in the Leader board but I understand how this works and it feels good :)</p>\n\n<p>2 - As Pedro below suggests, go to some old competitions ( Again ideally you should pick a problem that interests you), and slowly but steadily work through the data set. </p>\n\n<p>Pros and Cons </p>\n\n<p>Approach 1 - \nPros -  It will be a punch to the gut. It will make you hustle, work harder and make things happen in little ways for yourself. \nCons -  Competition mode can be addictive. It will mess with your mind :)</p>\n\n<p>Approach 2 - \nPros - Slow and steady approach without losing your sanity. You can of course work your way up to LIVE competitions \nCons - Nothing really. I guess one needs to have the internal drive to systematically work on different data sets</p>\n\n<p>Hope this helps \nRegards\nShanth </p>",
      "rawMarkdown": "JonW \n\nFirst up good decision. So Congrats. \nI was in your position about 4 months ago, so I know how you feel. \n\n2 Approaches \n\n1 - Jump headlong into a LIVE competition depending on your area of interest. I started with the Mercari Price Suggestion ( An NLP problem )challenge and I did not even know how to vectorize a sentence when I started this competition. However, going after a single problem gave me the focus to go out and try all kinds of approaches. The learning was incredible. It was frustrating, sleep depriving and down right discouraging as conversations flew thick and fast about things I simply could not follow. I hung in there, kept learning, crashed out of the competition without a rank but the learning was MASSIVE. I now have the confidence to approach and work on NLP problems. I am nowhere near winning or making a dent in the Leader board but I understand how this works and it feels good :)\n\n2 - As Pedro below suggests, go to some old competitions ( Again ideally you should pick a problem that interests you), and slowly but steadily work through the data set. \n\nPros and Cons \n\nApproach 1 - \nPros -  It will be a punch to the gut. It will make you hustle, work harder and make things happen in little ways for yourself. \nCons -  Competition mode can be addictive. It will mess with your mind :)\n\nApproach 2 - \nPros - Slow and steady approach without losing your sanity. You can of course work your way up to LIVE competitions \nCons - Nothing really. I guess one needs to have the internal drive to systematically work on different data sets\n\nHope this helps \nRegards\nShanth",
      "votes": null
    },
    {
      "id": "309113",
      "postDate": "04/04/2018 17:04:30",
      "content": "<p>Welcome! I was in the same boat as you but with R. I found Kaggle because we were learning R in my stats class. With 100% certainty, by simply examining others' work here and practicing it myself, this website turned me into a proficient user extremely quickly. So I think you are making the right decision by diving in :) </p>\n\n<p>My advice would be to take a look at other Kernels (especially ones structured like tutorials!) and in doing so, look at the codes and try to understand what's going on. Don't just robotically copy it. Really try to understand why it's coded the way it is and internalize it. </p>\n\n<p>Also, I would start small. I'm still hesitant to have a go at machine learning because I don't completely understand the concepts... but on a smaller scale where I can apply things I've already learned I do well because I can explain it. I even won a weekly award here by doing simple linear regression and that was the coolest thing ever! </p>\n\n<p>So basically... just do what you can here whether it's diving into a competition or just starting small on smaller datasets or previous competitions. Whatever you choose, doing is learning! Good luck!</p>",
      "rawMarkdown": "Welcome! I was in the same boat as you but with R. I found Kaggle because we were learning R in my stats class. With 100% certainty, by simply examining others' work here and practicing it myself, this website turned me into a proficient user extremely quickly. So I think you are making the right decision by diving in :) \n\nMy advice would be to take a look at other Kernels (especially ones structured like tutorials!) and in doing so, look at the codes and try to understand what's going on. Don't just robotically copy it. Really try to understand why it's coded the way it is and internalize it. \n\nAlso, I would start small. I'm still hesitant to have a go at machine learning because I don't completely understand the concepts... but on a smaller scale where I can apply things I've already learned I do well because I can explain it. I even won a weekly award here by doing simple linear regression and that was the coolest thing ever! \n\nSo basically... just do what you can here whether it's diving into a competition or just starting small on smaller datasets or previous competitions. Whatever you choose, doing is learning! Good luck!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 307323,
      "author_name": "pvlima",
      "author_url": "",
      "post_date": "04/01/2018 11:13:56",
      "content": "<p>Congratulations, you are in the right place to learn by doing :). My recommendation is to start with a smaller dataset (the ones in the Getting Started category). That will allow you to make many experiments fast. Then give yourself the goal of creating from scratch a simple model for the problem (linear or logistic regression for example) and make that work. Then go and try some other algorithms like random forest and gradient boosting. After being able to solve these small problems everything else will be quick. Enjoy.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 307425,
      "author_name": "asparuhhristov",
      "author_url": "",
      "post_date": "04/01/2018 15:35:40",
      "content": "<p>I agree that learning by doing is the best (at least for me), but nevertheless, I will combine this with studing existing kernels from other people out there. For example, try your best for 4-5 hours on one of the playgrounds or in a past competition and then look at 5-10 of the \"low-level\" kernels (search for keywords like EDA, or tutorial, or baseline). Study what you could do better (or more), maybe try to do it again or just go to the other competition. Repeat.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 308827,
      "author_name": "malten",
      "author_url": "",
      "post_date": "04/04/2018 07:06:53",
      "content": "<p>Personally i learn the most by writing my own code, as it forces you to think the problem through, otherwise it will not work out. Hard in the beginning but gets easier and easier, also very satisfying, when the program smoothly runs through. </p>\n\n<p>For a start (and basically all the way) begin with building a working pipeline, that is a program that loads in the data, calculates the features (start simple here), runs a predictive model, predicts the test data and writes a submission file that you can submit here (and get a realistic score &gt;0.93). Most kernels have such a pipeline, that you can use for inspiration, however i would recommend to write your own from scratch as this way you learn the most and start to think like a data scientist.</p>\n\n<p>I agree with the others though that this dataset is not great for starting, as the size requires additional considerations, if you dont have a lot of computational ressources. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 309072,
      "author_name": "shanth84",
      "author_url": "",
      "post_date": "04/04/2018 15:35:09",
      "content": "<p>@ JonW </p>\n\n<p>First up good decision. So Congrats. \nI was in your position about 4 months ago, so I know how you feel. </p>\n\n<p>2 Approaches </p>\n\n<p>1 - Jump headlong into a LIVE competition depending on your area of interest. I started with the Mercari Price Suggestion ( An NLP problem )challenge and I did not even know how to vectorize a sentence when I started this competition. However, going after a single problem gave me the focus to go out and try all kinds of approaches. The learning was incredible. It was frustrating, sleep depriving and down right discouraging as conversations flew thick and fast about things I simply could not follow. I hung in there, kept learning, crashed out of the competition without a rank but the learning was MASSIVE. I now have the confidence to approach and work on NLP problems. I am nowhere near winning or making a dent in the Leader board but I understand how this works and it feels good :)</p>\n\n<p>2 - As Pedro below suggests, go to some old competitions ( Again ideally you should pick a problem that interests you), and slowly but steadily work through the data set. </p>\n\n<p>Pros and Cons </p>\n\n<p>Approach 1 - \nPros -  It will be a punch to the gut. It will make you hustle, work harder and make things happen in little ways for yourself. \nCons -  Competition mode can be addictive. It will mess with your mind :)</p>\n\n<p>Approach 2 - \nPros - Slow and steady approach without losing your sanity. You can of course work your way up to LIVE competitions \nCons - Nothing really. I guess one needs to have the internal drive to systematically work on different data sets</p>\n\n<p>Hope this helps \nRegards\nShanth </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 309113,
      "author_name": "mistermichael",
      "author_url": "",
      "post_date": "04/04/2018 17:04:30",
      "content": "<p>Welcome! I was in the same boat as you but with R. I found Kaggle because we were learning R in my stats class. With 100% certainty, by simply examining others' work here and practicing it myself, this website turned me into a proficient user extremely quickly. So I think you are making the right decision by diving in :) </p>\n\n<p>My advice would be to take a look at other Kernels (especially ones structured like tutorials!) and in doing so, look at the codes and try to understand what's going on. Don't just robotically copy it. Really try to understand why it's coded the way it is and internalize it. </p>\n\n<p>Also, I would start small. I'm still hesitant to have a go at machine learning because I don't completely understand the concepts... but on a smaller scale where I can apply things I've already learned I do well because I can explain it. I even won a weekly award here by doing simple linear regression and that was the coolest thing ever! </p>\n\n<p>So basically... just do what you can here whether it's diving into a competition or just starting small on smaller datasets or previous competitions. Whatever you choose, doing is learning! Good luck!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "307301": "Hi,\nI feel a complete incompetent newbie... :-S\nStruggled my way through the first 3 courses of the Python for Everyone Specialization on Coursera, \nI saw someone mention Kaggle as a good place to practice with simple problems and cases. \nI feel quite overwhelmed by the Kernils and Tutorials voted high in the discussion board. I guess I need more foundation knowledge.... \nI am happy to learn more by watching videos or reading books, but is there a better way to really learn as an ABSOLUTE NEWBIE by DOING?\n\nAny encouragement and recommendations for someone who learns by doing, please?",
    "307323": "Congratulations, you are in the right place to learn by doing :). My recommendation is to start with a smaller dataset (the ones in the Getting Started category). That will allow you to make many experiments fast. Then give yourself the goal of creating from scratch a simple model for the problem (linear or logistic regression for example) and make that work. Then go and try some other algorithms like random forest and gradient boosting. After being able to solve these small problems everything else will be quick. Enjoy.",
    "307425": "I agree that learning by doing is the best (at least for me), but nevertheless, I will combine this with studing existing kernels from other people out there. For example, try your best for 4-5 hours on one of the playgrounds or in a past competition and then look at 5-10 of the \"low-level\" kernels (search for keywords like EDA, or tutorial, or baseline). Study what you could do better (or more), maybe try to do it again or just go to the other competition. Repeat.",
    "308827": "Personally i learn the most by writing my own code, as it forces you to think the problem through, otherwise it will not work out. Hard in the beginning but gets easier and easier, also very satisfying, when the program smoothly runs through. \n\n For a start (and basically all the way) begin with building a working pipeline, that is a program that loads in the data, calculates the features (start simple here), runs a predictive model, predicts the test data and writes a submission file that you can submit here (and get a realistic score &gt;0.93). Most kernels have such a pipeline, that you can use for inspiration, however i would recommend to write your own from scratch as this way you learn the most and start to think like a data scientist.\n\n I agree with the others though that this dataset is not great for starting, as the size requires additional considerations, if you dont have a lot of computational ressources.",
    "309072": "JonW \n\nFirst up good decision. So Congrats. \nI was in your position about 4 months ago, so I know how you feel. \n\n2 Approaches \n\n1 - Jump headlong into a LIVE competition depending on your area of interest. I started with the Mercari Price Suggestion ( An NLP problem )challenge and I did not even know how to vectorize a sentence when I started this competition. However, going after a single problem gave me the focus to go out and try all kinds of approaches. The learning was incredible. It was frustrating, sleep depriving and down right discouraging as conversations flew thick and fast about things I simply could not follow. I hung in there, kept learning, crashed out of the competition without a rank but the learning was MASSIVE. I now have the confidence to approach and work on NLP problems. I am nowhere near winning or making a dent in the Leader board but I understand how this works and it feels good :)\n\n2 - As Pedro below suggests, go to some old competitions ( Again ideally you should pick a problem that interests you), and slowly but steadily work through the data set. \n\nPros and Cons \n\nApproach 1 - \nPros -  It will be a punch to the gut. It will make you hustle, work harder and make things happen in little ways for yourself. \nCons -  Competition mode can be addictive. It will mess with your mind :)\n\nApproach 2 - \nPros - Slow and steady approach without losing your sanity. You can of course work your way up to LIVE competitions \nCons - Nothing really. I guess one needs to have the internal drive to systematically work on different data sets\n\nHope this helps \nRegards\nShanth",
    "309113": "Welcome! I was in the same boat as you but with R. I found Kaggle because we were learning R in my stats class. With 100% certainty, by simply examining others' work here and practicing it myself, this website turned me into a proficient user extremely quickly. So I think you are making the right decision by diving in :) \n\nMy advice would be to take a look at other Kernels (especially ones structured like tutorials!) and in doing so, look at the codes and try to understand what's going on. Don't just robotically copy it. Really try to understand why it's coded the way it is and internalize it. \n\nAlso, I would start small. I'm still hesitant to have a go at machine learning because I don't completely understand the concepts... but on a smaller scale where I can apply things I've already learned I do well because I can explain it. I even won a weekly award here by doing simple linear regression and that was the coolest thing ever! \n\nSo basically... just do what you can here whether it's diving into a competition or just starting small on smaller datasets or previous competitions. Whatever you choose, doing is learning! Good luck!"
  },
  "source": "meta"
}