{
  "id": 247316,
  "title": "Is this the \"Google Smartphone Data Preparation Challenge\"?",
  "url": "/competitions/google-smartphone-decimeter-challenge/discussion/247316",
  "author_name": "",
  "post_date": "2021-06-19T06:54:38.544575700Z",
  "votes": 21,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I don't know about anybody else but this seems more to me like the <strong>\"Google Smartphone Data Preparation Challenge\"</strong>; I have so far spent more than 90% of my time simply trying to create a decent pair of <code>train/test</code> datasets!</p>\n<p>It is definitely putting my ability to use pandas <code>merge</code> and <code>merge_asof</code> skills to the test. If nothing else it certainly gives me a new appreciation of how much work goes on behind the scenes at kaggle to provide us with the usual beautifully curated and ready-to-go <code>train/test</code> datasets that we are generally used to with most other competitions!</p>\n<p>All the best!<br>\ncarl </p>",
  "messages": [
    {
      "id": "1356677",
      "postDate": "06/19/2021 06:54:38",
      "content": "<p>I don't know about anybody else but this seems more to me like the <strong>\"Google Smartphone Data Preparation Challenge\"</strong>; I have so far spent more than 90% of my time simply trying to create a decent pair of <code>train/test</code> datasets!</p>\n<p>It is definitely putting my ability to use pandas <code>merge</code> and <code>merge_asof</code> skills to the test. If nothing else it certainly gives me a new appreciation of how much work goes on behind the scenes at kaggle to provide us with the usual beautifully curated and ready-to-go <code>train/test</code> datasets that we are generally used to with most other competitions!</p>\n<p>All the best!<br>\ncarl </p>",
      "rawMarkdown": "I don't know about anybody else but this seems more to me like the **\"Google Smartphone Data Preparation Challenge\"**; I have so far spent more than 90% of my time simply trying to create a decent pair of `train/test` datasets!\n\nIt is definitely putting my ability to use pandas `merge` and `merge_asof` skills to the test. If nothing else it certainly gives me a new appreciation of how much work goes on behind the scenes at kaggle to provide us with the usual beautifully curated and ready-to-go `train/test` datasets that we are generally used to with most other competitions!\n\nAll the best!\ncarl",
      "votes": null
    },
    {
      "id": "1358071",
      "postDate": "06/20/2021 07:09:46",
      "content": "<p>Data preparation is like a threshold for competition. For example, in my case, seeing the data, I don't know where to start.</p>",
      "rawMarkdown": "Data preparation is like a threshold for competition. For example, in my case, seeing the data, I don't know where to start.",
      "votes": null
    },
    {
      "id": "1358113",
      "postDate": "06/20/2021 07:48:27",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/yus002\" target=\"_blank\">@yus002</a> </p>\n<p>Threshold: indeed! One can certainly participate in this competition using the baseline <code>train</code> and <code>test</code> in conjunction with the ground truth files, as was wonderfully demonstrated right at the very outset of this competition by <a href=\"https://www.kaggle.com/emaerthin\" target=\"_blank\">@emaerthin</a> in the excellent notebook <a href=\"https://www.kaggle.com/emaerthin/demonstration-of-the-kalman-filter/data\" target=\"_blank\">\"<em>Demonstration of the Kalman filter</em>\"</a>. However, at some point one would really like to create an \"augmented\" dataset which incorporates the <code>derived</code> data along with the essential data found in the raw files. However, creating such an augmented dataset is where the fun begins!</p>\n<p>Let me be clear, although the title of this Topic may seem somewhat ironic, in some ways I think this is one of the most realistic competitions there is on kaggle, as it far better reflects the real day-to-day work of a data scientist. Kaggle primarily being a competitive machine learning site (although diversifying over the last few years, especially regarding datasets) hosts competitions, and a competition requires a leaderboard and thus a metric, and it is very hard to think up a \"competitive data preparation competition\" metric!</p>\n<p>In most tabular data competitions is entirely feasible to knock out a score that will be within 5% of the actual final winning score in just one sitting with, say, XGBoost (for example, one can score <code>0.78</code> on the Titanic in 5 minutes, but spend weeks getting <code>0.81</code> without overfitting) . The rest of the three months are composed of a struggle to squeeze out every last decimal point on the LB to eventually win. However, I suspect that the majority of employers will frown on such dedication of resources: most of the value will be in the data collection, preparation and feature engineering, a process which in its-self will contribute valuable business insights. </p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @yus002 \n\nThreshold: indeed! One can certainly participate in this competition using the baseline `train` and `test` in conjunction with the ground truth files, as was wonderfully demonstrated right at the very outset of this competition by @emaerthin in the excellent notebook [\"*Demonstration of the Kalman filter*\"](https://www.kaggle.com/emaerthin/demonstration-of-the-kalman-filter/data). However, at some point one would really like to create an \"augmented\" dataset which incorporates the `derived` data along with the essential data found in the raw files. However, creating such an augmented dataset is where the fun begins!\n\nLet me be clear, although the title of this Topic may seem somewhat ironic, in some ways I think this is one of the most realistic competitions there is on kaggle, as it far better reflects the real day-to-day work of a data scientist. Kaggle primarily being a competitive machine learning site (although diversifying over the last few years, especially regarding datasets) hosts competitions, and a competition requires a leaderboard and thus a metric, and it is very hard to think up a \"competitive data preparation competition\" metric!\n\nIn most tabular data competitions is entirely feasible to knock out a score that will be within 5% of the actual final winning score in just one sitting with, say, XGBoost (for example, one can score `0.78` on the Titanic in 5 minutes, but spend weeks getting `0.81` without overfitting) . The rest of the three months are composed of a struggle to squeeze out every last decimal point on the LB to eventually win. However, I suspect that the majority of employers will frown on such dedication of resources: most of the value will be in the data collection, preparation and feature engineering, a process which in its-self will contribute valuable business insights. \n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "1358121",
      "postDate": "06/20/2021 07:58:24",
      "content": "<p>wonderful comment</p>",
      "rawMarkdown": "wonderful comment",
      "votes": null
    },
    {
      "id": "1358253",
      "postDate": "06/20/2021 10:00:08",
      "content": "<p>Actually I still don't understand the data. 😅</p>",
      "rawMarkdown": "Actually I still don't understand the data. 😅",
      "votes": null
    },
    {
      "id": "1358555",
      "postDate": "06/20/2021 15:08:50",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/woosungyoon\" target=\"_blank\">@woosungyoon</a> </p>\n<p>For a great introduction to GPS, along with MATLAB code snippets, may I suggest reading through the excellent page <a href=\"https://www.telesens.co/2017/07/17/calculating-position-from-raw-gps-data/\" target=\"_blank\">\"<em>Calculating Position from Raw GPS Data</em>\"</a>.</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @woosungyoon \n\nFor a great introduction to GPS, along with MATLAB code snippets, may I suggest reading through the excellent page [\"*Calculating Position from Raw GPS Data*\"](https://www.telesens.co/2017/07/17/calculating-position-from-raw-gps-data/).\n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "1359117",
      "postDate": "06/21/2021 05:03:57",
      "content": "<p>On the one hand - this is very unlike most Kaggle competitions (with some exceptions like the indoor competition that took place a few weeks ago) so we do need to put a lot of effort into pre-processing,<br>\nBut this is a good way to work on what a Data Scientist works on every day, I think it's worthwhile to strengthen the pre-process muscles as well</p>",
      "rawMarkdown": "On the one hand - this is very unlike most Kaggle competitions (with some exceptions like the indoor competition that took place a few weeks ago) so we do need to put a lot of effort into pre-processing,\nBut this is a good way to work on what a Data Scientist works on every day, I think it's worthwhile to strengthen the pre-process muscles as well",
      "votes": null
    },
    {
      "id": "1388114",
      "postDate": "07/14/2021 16:53:15",
      "content": "<p>Real-world data science is 90% data cleansing and 10% modelling. This competition gives the taste of it</p>",
      "rawMarkdown": "Real-world data science is 90% data cleansing and 10% modelling. This competition gives the taste of it",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1358071,
      "author_name": "yus002",
      "author_url": "",
      "post_date": "06/20/2021 07:09:46",
      "content": "<p>Data preparation is like a threshold for competition. For example, in my case, seeing the data, I don't know where to start.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1358113,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "06/20/2021 07:48:27",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/yus002\" target=\"_blank\">@yus002</a> </p>\n<p>Threshold: indeed! One can certainly participate in this competition using the baseline <code>train</code> and <code>test</code> in conjunction with the ground truth files, as was wonderfully demonstrated right at the very outset of this competition by <a href=\"https://www.kaggle.com/emaerthin\" target=\"_blank\">@emaerthin</a> in the excellent notebook <a href=\"https://www.kaggle.com/emaerthin/demonstration-of-the-kalman-filter/data\" target=\"_blank\">\"<em>Demonstration of the Kalman filter</em>\"</a>. However, at some point one would really like to create an \"augmented\" dataset which incorporates the <code>derived</code> data along with the essential data found in the raw files. However, creating such an augmented dataset is where the fun begins!</p>\n<p>Let me be clear, although the title of this Topic may seem somewhat ironic, in some ways I think this is one of the most realistic competitions there is on kaggle, as it far better reflects the real day-to-day work of a data scientist. Kaggle primarily being a competitive machine learning site (although diversifying over the last few years, especially regarding datasets) hosts competitions, and a competition requires a leaderboard and thus a metric, and it is very hard to think up a \"competitive data preparation competition\" metric!</p>\n<p>In most tabular data competitions is entirely feasible to knock out a score that will be within 5% of the actual final winning score in just one sitting with, say, XGBoost (for example, one can score <code>0.78</code> on the Titanic in 5 minutes, but spend weeks getting <code>0.81</code> without overfitting) . The rest of the three months are composed of a struggle to squeeze out every last decimal point on the LB to eventually win. However, I suspect that the majority of employers will frown on such dedication of resources: most of the value will be in the data collection, preparation and feature engineering, a process which in its-self will contribute valuable business insights. </p>\n<p>All the best,<br>\ncarl</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1358121,
          "author_name": "yus002",
          "author_url": "",
          "post_date": "06/20/2021 07:58:24",
          "content": "<p>wonderful comment</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1358253,
      "author_name": "woosungyoon",
      "author_url": "",
      "post_date": "06/20/2021 10:00:08",
      "content": "<p>Actually I still don't understand the data. 😅</p>",
      "votes": null,
      "replies": [
        {
          "id": 1358555,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "06/20/2021 15:08:50",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/woosungyoon\" target=\"_blank\">@woosungyoon</a> </p>\n<p>For a great introduction to GPS, along with MATLAB code snippets, may I suggest reading through the excellent page <a href=\"https://www.telesens.co/2017/07/17/calculating-position-from-raw-gps-data/\" target=\"_blank\">\"<em>Calculating Position from Raw GPS Data</em>\"</a>.</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1359117,
      "author_name": "avivlevi815",
      "author_url": "",
      "post_date": "06/21/2021 05:03:57",
      "content": "<p>On the one hand - this is very unlike most Kaggle competitions (with some exceptions like the indoor competition that took place a few weeks ago) so we do need to put a lot of effort into pre-processing,<br>\nBut this is a good way to work on what a Data Scientist works on every day, I think it's worthwhile to strengthen the pre-process muscles as well</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1388114,
      "author_name": "shaz13",
      "author_url": "",
      "post_date": "07/14/2021 16:53:15",
      "content": "<p>Real-world data science is 90% data cleansing and 10% modelling. This competition gives the taste of it</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1356677": "I don't know about anybody else but this seems more to me like the **\"Google Smartphone Data Preparation Challenge\"**; I have so far spent more than 90% of my time simply trying to create a decent pair of `train/test` datasets!\n\nIt is definitely putting my ability to use pandas `merge` and `merge_asof` skills to the test. If nothing else it certainly gives me a new appreciation of how much work goes on behind the scenes at kaggle to provide us with the usual beautifully curated and ready-to-go `train/test` datasets that we are generally used to with most other competitions!\n\nAll the best!\ncarl",
    "1358071": "Data preparation is like a threshold for competition. For example, in my case, seeing the data, I don't know where to start.",
    "1358113": "Dear @yus002 \n\nThreshold: indeed! One can certainly participate in this competition using the baseline `train` and `test` in conjunction with the ground truth files, as was wonderfully demonstrated right at the very outset of this competition by @emaerthin in the excellent notebook [\"*Demonstration of the Kalman filter*\"](https://www.kaggle.com/emaerthin/demonstration-of-the-kalman-filter/data). However, at some point one would really like to create an \"augmented\" dataset which incorporates the `derived` data along with the essential data found in the raw files. However, creating such an augmented dataset is where the fun begins!\n\nLet me be clear, although the title of this Topic may seem somewhat ironic, in some ways I think this is one of the most realistic competitions there is on kaggle, as it far better reflects the real day-to-day work of a data scientist. Kaggle primarily being a competitive machine learning site (although diversifying over the last few years, especially regarding datasets) hosts competitions, and a competition requires a leaderboard and thus a metric, and it is very hard to think up a \"competitive data preparation competition\" metric!\n\nIn most tabular data competitions is entirely feasible to knock out a score that will be within 5% of the actual final winning score in just one sitting with, say, XGBoost (for example, one can score `0.78` on the Titanic in 5 minutes, but spend weeks getting `0.81` without overfitting) . The rest of the three months are composed of a struggle to squeeze out every last decimal point on the LB to eventually win. However, I suspect that the majority of employers will frown on such dedication of resources: most of the value will be in the data collection, preparation and feature engineering, a process which in its-self will contribute valuable business insights. \n\nAll the best,\ncarl",
    "1358121": "wonderful comment",
    "1358253": "Actually I still don't understand the data. 😅",
    "1358555": "Dear @woosungyoon \n\nFor a great introduction to GPS, along with MATLAB code snippets, may I suggest reading through the excellent page [\"*Calculating Position from Raw GPS Data*\"](https://www.telesens.co/2017/07/17/calculating-position-from-raw-gps-data/).\n\nAll the best,\ncarl",
    "1359117": "On the one hand - this is very unlike most Kaggle competitions (with some exceptions like the indoor competition that took place a few weeks ago) so we do need to put a lot of effort into pre-processing,\nBut this is a good way to work on what a Data Scientist works on every day, I think it's worthwhile to strengthen the pre-process muscles as well",
    "1388114": "Real-world data science is 90% data cleansing and 10% modelling. This competition gives the taste of it"
  },
  "source": "meta"
}