{
  "id": 244283,
  "title": "How to Approach This Competition",
  "url": "/competitions/google-smartphone-decimeter-challenge/discussion/244283",
  "author_name": "",
  "post_date": "2021-06-06T01:20:24.420135900Z",
  "votes": 40,
  "comment_count": 4,
  "views": 0,
  "content": "<p>GPS data isn’t the simplest to work with, so I hope some of these tips help you out.</p>\n<p>1 - Understand the format of this dataset.<br>\n<a href=\"https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/241039\" target=\"_blank\">Linked is a useful post to help with that.</a></p>\n<p>2 - Less Machine Learning<br>\nMachine learning algorithms (mostly referring to gradient boosted machines and neural networks) can definitely be useful in just about any kaggle challenge. That being said, when working with this data, it is important to note that the baseline_locations_test file has a latDeg and lngDeg column that when placed into the submission file already produces a pretty good score of around 7.19. For this reason, other post processing techniques may produce more valuable results than machine learning. </p>\n<p>3 - Correct outliers<br>\nSince we are working with data in time series paths, correcting outliers can be especially useful and should probably be one of your first steps. <a href=\"https://www.kaggle.com/dehokanta/baseline-post-processing-by-outlier-correction\" target=\"_blank\">Linked is a good example to start with.</a></p>\n<p>4 - Smoothing Data &amp; Kalman Filter<br>\nData smoothing uses an algorithm to remove noise from a data set. This can be very useful when working with gps data. A really powerful method is the kalman filter (<a href=\"https://www.kaggle.com/emaerthin/demonstration-of-the-kalman-filter\" target=\"_blank\">example linked</a>). I would highly recommend trying many types of data smoothing as there are multiple variants of the kalman filter and other forms of smoothening besides kalman.</p>\n<p>5 - Post Processing<br>\nThere are several other types of post processing that you can use besides data smoothing. Reading up on post processing for geospatial data may help give you an edge.</p>\n<p>6 - Hyperparameter Tune<br>\nTuning your post processing methods and models can help you squeeze the best results out of the data. You can use the train baseline and ground truths to help you tune these models. (<a href=\"https://www.kaggle.com/tqa236/kalman-filter-hyperparameter-search-with-bo\" target=\"_blank\">Linked is an example</a>).</p>\n<p>7 - Explore all the data<br>\nWhile I haven’t fully investigated and used the whole dataset, I am sure there are good ways to use the gnss and derived files. (<a href=\"https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/238590\" target=\"_blank\">check out this post</a>)</p>",
  "messages": [
    {
      "id": "1337871",
      "postDate": "06/06/2021 01:20:24",
      "content": "<p>GPS data isn’t the simplest to work with, so I hope some of these tips help you out.</p>\n<p>1 - Understand the format of this dataset.<br>\n<a href=\"https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/241039\" target=\"_blank\">Linked is a useful post to help with that.</a></p>\n<p>2 - Less Machine Learning<br>\nMachine learning algorithms (mostly referring to gradient boosted machines and neural networks) can definitely be useful in just about any kaggle challenge. That being said, when working with this data, it is important to note that the baseline_locations_test file has a latDeg and lngDeg column that when placed into the submission file already produces a pretty good score of around 7.19. For this reason, other post processing techniques may produce more valuable results than machine learning. </p>\n<p>3 - Correct outliers<br>\nSince we are working with data in time series paths, correcting outliers can be especially useful and should probably be one of your first steps. <a href=\"https://www.kaggle.com/dehokanta/baseline-post-processing-by-outlier-correction\" target=\"_blank\">Linked is a good example to start with.</a></p>\n<p>4 - Smoothing Data &amp; Kalman Filter<br>\nData smoothing uses an algorithm to remove noise from a data set. This can be very useful when working with gps data. A really powerful method is the kalman filter (<a href=\"https://www.kaggle.com/emaerthin/demonstration-of-the-kalman-filter\" target=\"_blank\">example linked</a>). I would highly recommend trying many types of data smoothing as there are multiple variants of the kalman filter and other forms of smoothening besides kalman.</p>\n<p>5 - Post Processing<br>\nThere are several other types of post processing that you can use besides data smoothing. Reading up on post processing for geospatial data may help give you an edge.</p>\n<p>6 - Hyperparameter Tune<br>\nTuning your post processing methods and models can help you squeeze the best results out of the data. You can use the train baseline and ground truths to help you tune these models. (<a href=\"https://www.kaggle.com/tqa236/kalman-filter-hyperparameter-search-with-bo\" target=\"_blank\">Linked is an example</a>).</p>\n<p>7 - Explore all the data<br>\nWhile I haven’t fully investigated and used the whole dataset, I am sure there are good ways to use the gnss and derived files. (<a href=\"https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/238590\" target=\"_blank\">check out this post</a>)</p>",
      "rawMarkdown": "GPS data isn’t the simplest to work with, so I hope some of these tips help you out.\n\n1 - Understand the format of this dataset.\n[Linked is a useful post to help with that.](https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/241039)\n\n2 - Less Machine Learning\nMachine learning algorithms (mostly referring to gradient boosted machines and neural networks) can definitely be useful in just about any kaggle challenge. That being said, when working with this data, it is important to note that the baseline_locations_test file has a latDeg and lngDeg column that when placed into the submission file already produces a pretty good score of around 7.19. For this reason, other post processing techniques may produce more valuable results than machine learning. \n\n3 - Correct outliers\nSince we are working with data in time series paths, correcting outliers can be especially useful and should probably be one of your first steps. [Linked is a good example to start with.](https://www.kaggle.com/dehokanta/baseline-post-processing-by-outlier-correction)\n\n4 - Smoothing Data & Kalman Filter\nData smoothing uses an algorithm to remove noise from a data set. This can be very useful when working with gps data. A really powerful method is the kalman filter ([example linked](https://www.kaggle.com/emaerthin/demonstration-of-the-kalman-filter)). I would highly recommend trying many types of data smoothing as there are multiple variants of the kalman filter and other forms of smoothening besides kalman.\n\n5 - Post Processing\nThere are several other types of post processing that you can use besides data smoothing. Reading up on post processing for geospatial data may help give you an edge.\n\n6 - Hyperparameter Tune\nTuning your post processing methods and models can help you squeeze the best results out of the data. You can use the train baseline and ground truths to help you tune these models. ([Linked is an example](https://www.kaggle.com/tqa236/kalman-filter-hyperparameter-search-with-bo)).\n\n7 - Explore all the data\nWhile I haven’t fully investigated and used the whole dataset, I am sure there are good ways to use the gnss and derived files. ([check out this post](https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/238590))",
      "votes": null
    },
    {
      "id": "1337917",
      "postDate": "06/06/2021 03:02:19",
      "content": "<p>Thanks for the great explanation. <br>\nI am now trying to study the content of the competition.<br>\nThis article has been helpful. </p>\n<p>I think this competition is dealing with the problem of location uncertainty. </p>",
      "rawMarkdown": "Thanks for the great explanation. \nI am now trying to study the content of the competition.\nThis article has been helpful. \n\nI think this competition is dealing with the problem of location uncertainty.",
      "votes": null
    },
    {
      "id": "1353416",
      "postDate": "06/17/2021 05:11:11",
      "content": "<p>Thanks Ravi</p>",
      "rawMarkdown": "Thanks Ravi",
      "votes": null
    },
    {
      "id": "1364079",
      "postDate": "06/24/2021 15:42:02",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">@ravishah1</a>, very useful information!</p>",
      "rawMarkdown": "Thanks for sharing @ravishah1, very useful information!",
      "votes": null
    },
    {
      "id": "1391728",
      "postDate": "07/18/2021 02:39:48",
      "content": "<p>Thanks Ravi - this is helpful.</p>",
      "rawMarkdown": "Thanks Ravi - this is helpful.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1337917,
      "author_name": "woosungyoon",
      "author_url": "",
      "post_date": "06/06/2021 03:02:19",
      "content": "<p>Thanks for the great explanation. <br>\nI am now trying to study the content of the competition.<br>\nThis article has been helpful. </p>\n<p>I think this competition is dealing with the problem of location uncertainty. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1353416,
      "author_name": "rakshittra",
      "author_url": "",
      "post_date": "06/17/2021 05:11:11",
      "content": "<p>Thanks Ravi</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1364079,
      "author_name": "saurabhbagchi",
      "author_url": "",
      "post_date": "06/24/2021 15:42:02",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">@ravishah1</a>, very useful information!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1391728,
      "author_name": "patungire",
      "author_url": "",
      "post_date": "07/18/2021 02:39:48",
      "content": "<p>Thanks Ravi - this is helpful.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1337871": "GPS data isn’t the simplest to work with, so I hope some of these tips help you out.\n\n1 - Understand the format of this dataset.\n[Linked is a useful post to help with that.](https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/241039)\n\n2 - Less Machine Learning\nMachine learning algorithms (mostly referring to gradient boosted machines and neural networks) can definitely be useful in just about any kaggle challenge. That being said, when working with this data, it is important to note that the baseline_locations_test file has a latDeg and lngDeg column that when placed into the submission file already produces a pretty good score of around 7.19. For this reason, other post processing techniques may produce more valuable results than machine learning. \n\n3 - Correct outliers\nSince we are working with data in time series paths, correcting outliers can be especially useful and should probably be one of your first steps. [Linked is a good example to start with.](https://www.kaggle.com/dehokanta/baseline-post-processing-by-outlier-correction)\n\n4 - Smoothing Data & Kalman Filter\nData smoothing uses an algorithm to remove noise from a data set. This can be very useful when working with gps data. A really powerful method is the kalman filter ([example linked](https://www.kaggle.com/emaerthin/demonstration-of-the-kalman-filter)). I would highly recommend trying many types of data smoothing as there are multiple variants of the kalman filter and other forms of smoothening besides kalman.\n\n5 - Post Processing\nThere are several other types of post processing that you can use besides data smoothing. Reading up on post processing for geospatial data may help give you an edge.\n\n6 - Hyperparameter Tune\nTuning your post processing methods and models can help you squeeze the best results out of the data. You can use the train baseline and ground truths to help you tune these models. ([Linked is an example](https://www.kaggle.com/tqa236/kalman-filter-hyperparameter-search-with-bo)).\n\n7 - Explore all the data\nWhile I haven’t fully investigated and used the whole dataset, I am sure there are good ways to use the gnss and derived files. ([check out this post](https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/238590))",
    "1337917": "Thanks for the great explanation. \nI am now trying to study the content of the competition.\nThis article has been helpful. \n\nI think this competition is dealing with the problem of location uncertainty.",
    "1353416": "Thanks Ravi",
    "1364079": "Thanks for sharing @ravishah1, very useful information!",
    "1391728": "Thanks Ravi - this is helpful."
  },
  "source": "meta"
}