{
  "id": 206042,
  "title": "How this competition affected my learning journey",
  "url": "/competitions/riiid-test-answer-prediction/discussion/206042",
  "author_name": "",
  "post_date": "2020-12-22T22:45:37.818597200Z",
  "votes": 7,
  "comment_count": 4,
  "views": 0,
  "content": "<p>This was the first featured competition I entered on Kaggle. I signed up two months ago fresh out of completing a Udacity course and entering the top 10% in both the House Prices and Titanic competitions. I felt ready to tackle the challenges of the competition and was keen to get stuck in.</p>\n<p>Then I looked at the data. Fair to say this is an unstructured dataset (am I right?). As such, I never submitted any predictions as <strong>I couldn't get my head around the data wrangling aspect of the competition.</strong></p>\n<p>And now I feel like a complete noob, despite all the things I've learned over last few months. I'm looking for some <strong>advice</strong> on how to better upskill in data wrangling as I feel it is a weak point in my machine learning knowledge base. I've read that the best data scientists are well versed in all disciplines i.e. they know how to wrangle data, build models and communicate the results. I am struggling with the first one.</p>\n<p>Any comments welcome </p>\n<p>Thanks</p>\n<p>Jon</p>",
  "messages": [
    {
      "id": "1123099",
      "postDate": "12/22/2020 22:45:37",
      "content": "<p>This was the first featured competition I entered on Kaggle. I signed up two months ago fresh out of completing a Udacity course and entering the top 10% in both the House Prices and Titanic competitions. I felt ready to tackle the challenges of the competition and was keen to get stuck in.</p>\n<p>Then I looked at the data. Fair to say this is an unstructured dataset (am I right?). As such, I never submitted any predictions as <strong>I couldn't get my head around the data wrangling aspect of the competition.</strong></p>\n<p>And now I feel like a complete noob, despite all the things I've learned over last few months. I'm looking for some <strong>advice</strong> on how to better upskill in data wrangling as I feel it is a weak point in my machine learning knowledge base. I've read that the best data scientists are well versed in all disciplines i.e. they know how to wrangle data, build models and communicate the results. I am struggling with the first one.</p>\n<p>Any comments welcome </p>\n<p>Thanks</p>\n<p>Jon</p>",
      "rawMarkdown": "This was the first featured competition I entered on Kaggle. I signed up two months ago fresh out of completing a Udacity course and entering the top 10% in both the House Prices and Titanic competitions. I felt ready to tackle the challenges of the competition and was keen to get stuck in.\n\nThen I looked at the data. Fair to say this is an unstructured dataset (am I right?). As such, I never submitted any predictions as **I couldn't get my head around the data wrangling aspect of the competition.**\n\nAnd now I feel like a complete noob, despite all the things I've learned over last few months. I'm looking for some **advice** on how to better upskill in data wrangling as I feel it is a weak point in my machine learning knowledge base. I've read that the best data scientists are well versed in all disciplines i.e. they know how to wrangle data, build models and communicate the results. I am struggling with the first one.\n\nAny comments welcome \n\nThanks\n\nJon",
      "votes": null
    },
    {
      "id": "1123111",
      "postDate": "12/22/2020 23:28:38",
      "content": "<p>Hi Jonathan !</p>\n<p>I would say it is perfectly normal at first to be completly lost when you are facing a brand new problems. </p>\n<p>Competitions such as Titanic have a dataset that is already all setup, and you just need to plug your favorite ML algorithm in it.</p>\n<p>In the case of this competition, one part of the job (if you go with standard ML models), is to generate your own features for the problem. </p>\n<p>Here are some of my advices: </p>\n<ul>\n<li>First, start small. To explore and \"take the temperature\", you can play with a subsample of the main dataset, lets say, for example, the first 100 000 rows. You can then increase when you will feel more comfortable manipulating the data.<br>\nA simple way to do it if you use pandas is to go with <code>pd.read_csv(... , nrows = 100000)</code> so you will avoid frustrating OOM.</li>\n<li>It is important to visualise the data and make yourself some statistics. There is some very nice EDA on public notebook, but try to visualise the information by yourself. What questions do you ask yourself when you see this data (for exemple: what are the most difficult questions ? How a user is progressing over time), and by what type of queries you could answer them ?</li>\n<li>Try to improve your skill in visualisation, you can start with matplotlib, and go further with plotly for example, that I personnaly prefer.</li>\n<li>Try to build your first set of features and start with simple ones: for a given line in the dataset, what information do you need to make a prediction ? You can start with very generic things such as general difficulty of a question, total number of questions seen by each id, the average reaction time, etc…</li>\n<li>Try to fit a model with those features, and keep the score as a baseline for futur improvements.<br>\n-Don't hesitate to look at public notebook to see what others are doing. One BIG thing about it: never copy/paste a solution or the code for a feature, but rather recode it by yourself. It will be much usefull for you and that's how you will progress. </li>\n<li>And a last but easy one: don't give up and don't hesitate to ask questions if you have any doubts, many people here will be happy to help.</li>\n</ul>\n<p>edit: Also another advice: you need to learn advance tricks with numpy and pandas (pivot tables, cumulative sum with resets, etc…)</p>",
      "rawMarkdown": "Hi Jonathan !\n\nI would say it is perfectly normal at first to be completly lost when you are facing a brand new problems. \n\nCompetitions such as Titanic have a dataset that is already all setup, and you just need to plug your favorite ML algorithm in it.\n\nIn the case of this competition, one part of the job (if you go with standard ML models), is to generate your own features for the problem. \n\nHere are some of my advices: \n- First, start small. To explore and \"take the temperature\", you can play with a subsample of the main dataset, lets say, for example, the first 100 000 rows. You can then increase when you will feel more comfortable manipulating the data.\nA simple way to do it if you use pandas is to go with `pd.read_csv(... , nrows = 100000)` so you will avoid frustrating OOM.\n- It is important to visualise the data and make yourself some statistics. There is some very nice EDA on public notebook, but try to visualise the information by yourself. What questions do you ask yourself when you see this data (for exemple: what are the most difficult questions ? How a user is progressing over time), and by what type of queries you could answer them ?\n- Try to improve your skill in visualisation, you can start with matplotlib, and go further with plotly for example, that I personnaly prefer.\n- Try to build your first set of features and start with simple ones: for a given line in the dataset, what information do you need to make a prediction ? You can start with very generic things such as general difficulty of a question, total number of questions seen by each id, the average reaction time, etc...\n- Try to fit a model with those features, and keep the score as a baseline for futur improvements.\n-Don't hesitate to look at public notebook to see what others are doing. One BIG thing about it: never copy/paste a solution or the code for a feature, but rather recode it by yourself. It will be much usefull for you and that's how you will progress. \n- And a last but easy one: don't give up and don't hesitate to ask questions if you have any doubts, many people here will be happy to help.\n\nedit: Also another advice: you need to learn advance tricks with numpy and pandas (pivot tables, cumulative sum with resets, etc...)",
      "votes": null
    },
    {
      "id": "1123142",
      "postDate": "12/23/2020 00:28:35",
      "content": "<p>Well, I think better to increase the difficulty level gradually before you get bad mood :)<br>\nThis competition is way way harder than House Prices and Titanic, because </p>\n<ul>\n<li>it's a time-series competition</li>\n<li>large dataset</li>\n<li>simulation api</li>\n<li>etc</li>\n</ul>\n<p>But, there are more similar competitions for your previous ones, so you can try out yourself on them too.<br>\nFor example if you like image classification, there is the ongoing Cassava Leaf Disease Classification: <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification</a></p>\n<p>I mean similar in style, not too big dataset in GB, more straightforward predictions.<br>\nIf you try that you can still do this competition and any/every other ongoing competitions - one/more will be your favorite(s) ;)<br>\nSo keep going ^^</p>\n<p>Here you can get them all.<br>\n<a href=\"https://www.kaggle.com/competitions\" target=\"_blank\">https://www.kaggle.com/competitions</a></p>",
      "rawMarkdown": "Well, I think better to increase the difficulty level gradually before you get bad mood :)\nThis competition is way way harder than House Prices and Titanic, because \n  - it's a time-series competition\n  - large dataset\n  - simulation api\n  - etc\n\nBut, there are more similar competitions for your previous ones, so you can try out yourself on them too.\nFor example if you like image classification, there is the ongoing Cassava Leaf Disease Classification: https://www.kaggle.com/c/cassava-leaf-disease-classification\n\nI mean similar in style, not too big dataset in GB, more straightforward predictions.\nIf you try that you can still do this competition and any/every other ongoing competitions - one/more will be your favorite(s) ;)\nSo keep going ^^\n\nHere you can get them all.\nhttps://www.kaggle.com/competitions",
      "votes": null
    },
    {
      "id": "1123604",
      "postDate": "12/23/2020 11:14:03",
      "content": "<p>This is not an advise. But this is what I did in past and also doing right now. I study various notebooks from past competitions. Like there are various EDA notebooks in every competitions. So I study various methods of EDA in order to get the sense of data. TO do data wrangling , u need to first understand data. So first understand your data deeply. Then proceed. <a href=\"https://www.kaggle.com/bowdenjr\" target=\"_blank\">@bowdenjr</a> </p>",
      "rawMarkdown": "This is not an advise. But this is what I did in past and also doing right now. I study various notebooks from past competitions. Like there are various EDA notebooks in every competitions. So I study various methods of EDA in order to get the sense of data. TO do data wrangling , u need to first understand data. So first understand your data deeply. Then proceed. @bowdenjr",
      "votes": null
    },
    {
      "id": "1123716",
      "postDate": "12/23/2020 12:59:46",
      "content": "<p>Thanks, this is really great advice. I think the \"start small\" thing is definitely on point - the problem I had was knowing where to even start with such a large and confusing dataset, but as you say if I go with 10,000 rows or something and start with exploring one or two features instead of being overwhelmed (which is definitely what was happening). </p>\n<p>I don't do much EDA because I'd rather be fitting the model and yet I know that we can only fit good models if we put in the work to understand the data. So I am my own enemy here I think. It's funny because I know that data understanding and cleaning makes up like 80% of our work, but I guess it's hard when everything is in a weird format.</p>",
      "rawMarkdown": "Thanks, this is really great advice. I think the \"start small\" thing is definitely on point - the problem I had was knowing where to even start with such a large and confusing dataset, but as you say if I go with 10,000 rows or something and start with exploring one or two features instead of being overwhelmed (which is definitely what was happening). \n\nI don't do much EDA because I'd rather be fitting the model and yet I know that we can only fit good models if we put in the work to understand the data. So I am my own enemy here I think. It's funny because I know that data understanding and cleaning makes up like 80% of our work, but I guess it's hard when everything is in a weird format.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1123111,
      "author_name": "bowaka",
      "author_url": "",
      "post_date": "12/22/2020 23:28:38",
      "content": "<p>Hi Jonathan !</p>\n<p>I would say it is perfectly normal at first to be completly lost when you are facing a brand new problems. </p>\n<p>Competitions such as Titanic have a dataset that is already all setup, and you just need to plug your favorite ML algorithm in it.</p>\n<p>In the case of this competition, one part of the job (if you go with standard ML models), is to generate your own features for the problem. </p>\n<p>Here are some of my advices: </p>\n<ul>\n<li>First, start small. To explore and \"take the temperature\", you can play with a subsample of the main dataset, lets say, for example, the first 100 000 rows. You can then increase when you will feel more comfortable manipulating the data.<br>\nA simple way to do it if you use pandas is to go with <code>pd.read_csv(... , nrows = 100000)</code> so you will avoid frustrating OOM.</li>\n<li>It is important to visualise the data and make yourself some statistics. There is some very nice EDA on public notebook, but try to visualise the information by yourself. What questions do you ask yourself when you see this data (for exemple: what are the most difficult questions ? How a user is progressing over time), and by what type of queries you could answer them ?</li>\n<li>Try to improve your skill in visualisation, you can start with matplotlib, and go further with plotly for example, that I personnaly prefer.</li>\n<li>Try to build your first set of features and start with simple ones: for a given line in the dataset, what information do you need to make a prediction ? You can start with very generic things such as general difficulty of a question, total number of questions seen by each id, the average reaction time, etc…</li>\n<li>Try to fit a model with those features, and keep the score as a baseline for futur improvements.<br>\n-Don't hesitate to look at public notebook to see what others are doing. One BIG thing about it: never copy/paste a solution or the code for a feature, but rather recode it by yourself. It will be much usefull for you and that's how you will progress. </li>\n<li>And a last but easy one: don't give up and don't hesitate to ask questions if you have any doubts, many people here will be happy to help.</li>\n</ul>\n<p>edit: Also another advice: you need to learn advance tricks with numpy and pandas (pivot tables, cumulative sum with resets, etc…)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1123716,
          "author_name": "bowdenjr",
          "author_url": "",
          "post_date": "12/23/2020 12:59:46",
          "content": "<p>Thanks, this is really great advice. I think the \"start small\" thing is definitely on point - the problem I had was knowing where to even start with such a large and confusing dataset, but as you say if I go with 10,000 rows or something and start with exploring one or two features instead of being overwhelmed (which is definitely what was happening). </p>\n<p>I don't do much EDA because I'd rather be fitting the model and yet I know that we can only fit good models if we put in the work to understand the data. So I am my own enemy here I think. It's funny because I know that data understanding and cleaning makes up like 80% of our work, but I guess it's hard when everything is in a weird format.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1123142,
      "author_name": "killimi",
      "author_url": "",
      "post_date": "12/23/2020 00:28:35",
      "content": "<p>Well, I think better to increase the difficulty level gradually before you get bad mood :)<br>\nThis competition is way way harder than House Prices and Titanic, because </p>\n<ul>\n<li>it's a time-series competition</li>\n<li>large dataset</li>\n<li>simulation api</li>\n<li>etc</li>\n</ul>\n<p>But, there are more similar competitions for your previous ones, so you can try out yourself on them too.<br>\nFor example if you like image classification, there is the ongoing Cassava Leaf Disease Classification: <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification</a></p>\n<p>I mean similar in style, not too big dataset in GB, more straightforward predictions.<br>\nIf you try that you can still do this competition and any/every other ongoing competitions - one/more will be your favorite(s) ;)<br>\nSo keep going ^^</p>\n<p>Here you can get them all.<br>\n<a href=\"https://www.kaggle.com/competitions\" target=\"_blank\">https://www.kaggle.com/competitions</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1123604,
      "author_name": "saurabhshahane",
      "author_url": "",
      "post_date": "12/23/2020 11:14:03",
      "content": "<p>This is not an advise. But this is what I did in past and also doing right now. I study various notebooks from past competitions. Like there are various EDA notebooks in every competitions. So I study various methods of EDA in order to get the sense of data. TO do data wrangling , u need to first understand data. So first understand your data deeply. Then proceed. <a href=\"https://www.kaggle.com/bowdenjr\" target=\"_blank\">@bowdenjr</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1123099": "This was the first featured competition I entered on Kaggle. I signed up two months ago fresh out of completing a Udacity course and entering the top 10% in both the House Prices and Titanic competitions. I felt ready to tackle the challenges of the competition and was keen to get stuck in.\n\nThen I looked at the data. Fair to say this is an unstructured dataset (am I right?). As such, I never submitted any predictions as **I couldn't get my head around the data wrangling aspect of the competition.**\n\nAnd now I feel like a complete noob, despite all the things I've learned over last few months. I'm looking for some **advice** on how to better upskill in data wrangling as I feel it is a weak point in my machine learning knowledge base. I've read that the best data scientists are well versed in all disciplines i.e. they know how to wrangle data, build models and communicate the results. I am struggling with the first one.\n\nAny comments welcome \n\nThanks\n\nJon",
    "1123111": "Hi Jonathan !\n\nI would say it is perfectly normal at first to be completly lost when you are facing a brand new problems. \n\nCompetitions such as Titanic have a dataset that is already all setup, and you just need to plug your favorite ML algorithm in it.\n\nIn the case of this competition, one part of the job (if you go with standard ML models), is to generate your own features for the problem. \n\nHere are some of my advices: \n- First, start small. To explore and \"take the temperature\", you can play with a subsample of the main dataset, lets say, for example, the first 100 000 rows. You can then increase when you will feel more comfortable manipulating the data.\nA simple way to do it if you use pandas is to go with `pd.read_csv(... , nrows = 100000)` so you will avoid frustrating OOM.\n- It is important to visualise the data and make yourself some statistics. There is some very nice EDA on public notebook, but try to visualise the information by yourself. What questions do you ask yourself when you see this data (for exemple: what are the most difficult questions ? How a user is progressing over time), and by what type of queries you could answer them ?\n- Try to improve your skill in visualisation, you can start with matplotlib, and go further with plotly for example, that I personnaly prefer.\n- Try to build your first set of features and start with simple ones: for a given line in the dataset, what information do you need to make a prediction ? You can start with very generic things such as general difficulty of a question, total number of questions seen by each id, the average reaction time, etc...\n- Try to fit a model with those features, and keep the score as a baseline for futur improvements.\n-Don't hesitate to look at public notebook to see what others are doing. One BIG thing about it: never copy/paste a solution or the code for a feature, but rather recode it by yourself. It will be much usefull for you and that's how you will progress. \n- And a last but easy one: don't give up and don't hesitate to ask questions if you have any doubts, many people here will be happy to help.\n\nedit: Also another advice: you need to learn advance tricks with numpy and pandas (pivot tables, cumulative sum with resets, etc...)",
    "1123142": "Well, I think better to increase the difficulty level gradually before you get bad mood :)\nThis competition is way way harder than House Prices and Titanic, because \n  - it's a time-series competition\n  - large dataset\n  - simulation api\n  - etc\n\nBut, there are more similar competitions for your previous ones, so you can try out yourself on them too.\nFor example if you like image classification, there is the ongoing Cassava Leaf Disease Classification: https://www.kaggle.com/c/cassava-leaf-disease-classification\n\nI mean similar in style, not too big dataset in GB, more straightforward predictions.\nIf you try that you can still do this competition and any/every other ongoing competitions - one/more will be your favorite(s) ;)\nSo keep going ^^\n\nHere you can get them all.\nhttps://www.kaggle.com/competitions",
    "1123604": "This is not an advise. But this is what I did in past and also doing right now. I study various notebooks from past competitions. Like there are various EDA notebooks in every competitions. So I study various methods of EDA in order to get the sense of data. TO do data wrangling , u need to first understand data. So first understand your data deeply. Then proceed. @bowdenjr",
    "1123716": "Thanks, this is really great advice. I think the \"start small\" thing is definitely on point - the problem I had was knowing where to even start with such a large and confusing dataset, but as you say if I go with 10,000 rows or something and start with exploring one or two features instead of being overwhelmed (which is definitely what was happening). \n\nI don't do much EDA because I'd rather be fitting the model and yet I know that we can only fit good models if we put in the work to understand the data. So I am my own enemy here I think. It's funny because I know that data understanding and cleaning makes up like 80% of our work, but I guess it's hard when everything is in a weird format."
  },
  "source": "meta"
}