{
  "id": 323548,
  "title": "How to Approach this Competition",
  "url": "/competitions/smartphone-decimeter-2022/discussion/323548",
  "author_name": "Ravi Shah",
  "post_date": "2022-05-07T04:09:12.030000",
  "votes": 65,
  "comment_count": 11,
  "views": 0,
  "content": "<p><strong>1 - Understand the format of this dataset, objective, and evaluation</strong> <br>\nI recommend getting familiar with the structure of the data and going through the <a href=\"https://www.kaggle.com/competitions/smartphone-decimeter-2022/data\" target=\"_blank\">competition data tab</a> </p>\n<p>The goal of this competition is to predict the exact latitude and longitude coordinates of a car based on GNSS data coming from cell phones. The submission file will then be scored using the mean of the 50th and 95th percentile distance errors (<a href=\"https://www.kaggle.com/competitions/smartphone-decimeter-2022/overview/evaluation\" target=\"_blank\">see more details</a>)</p>\n<p><strong>2 - Make a baseline</strong><br>\nSince the baseline for this year's competition is in ECEF (Earth-Centered Earth-Fixed) coordinate system, the coordinate system must be converted to BLH to get the latitude and longitude submission baseline.<br>\n<a href=\"https://www.kaggle.com/saitodevel01\" target=\"_blank\">@saitodevel01</a> has published <a href=\"https://www.kaggle.com/code/saitodevel01/gsdc2-baseline-submission\" target=\"_blank\">an excellent notebook</a> demonstrating how to get to this baseline. <br>\nThe baseline this year seems to be a public score of 4.87</p>\n<p>After you have this baseline, you basically have the path the car took, but the signal is noisy. Now you must correct the baseline and find the true path.</p>\n<p><strong>3 - Correct Outliers</strong><br>\nSince we are working with data in time series paths, correcting outliers can be especially useful and should probably be one of your first steps. <a href=\"https://www.kaggle.com/code/ravishah1/gsdc2-savgol-filter-outlier-removal-with-bo\" target=\"_blank\">In my notebook</a>, I demonstrate a popular method of outlier correction I saw used in last year's competition.</p>\n<p><strong>4 - Smooth Baseline</strong><br>\nData smoothing uses an algorithm to remove noise from a data set. This can be very useful when working with gnss data. <a href=\"https://www.kaggle.com/code/ravishah1/gsdc2-savgol-filter-outlier-removal-with-bo\" target=\"_blank\">In my notebook</a>, I also demonstrate how to use the savgol filter to smooth the baseline. </p>\n<p>There are also many other forms of data smoothing you can try such as the Kalman filter which was popular in <a href=\"https://www.kaggle.com/code/emaerthin/demonstration-of-the-kalman-filter\" target=\"_blank\">public notebooks</a> such as this one by <a href=\"https://www.kaggle.com/emaerthin\" target=\"_blank\">@emaerthin</a> in last year's competition. </p>\n<p><strong>5 - Post Processing</strong><br>\nThere are several other types of post processing that you can use besides data smoothing. A lot of ideas such as snap to grid were used last year. Reading up on post processing for geospatial data may help give you an edge.<br>\nTo learn more about gnss data check out the resources in <a href=\"https://www.kaggle.com/chris62\" target=\"_blank\">@chris62</a> <a href=\"https://www.kaggle.com/competitions/smartphone-decimeter-2022/discussion/323057\" target=\"_blank\">discussion</a></p>\n<p><strong>6 - Machine Learning</strong><br>\nMachine learning may be a more challenging approach in this competition and post processing may be a simpler and more rewarding approach to start with. Nonetheless, using machine learning effectively may give you an edge in this competition. <br>\nSee how <a href=\"https://www.kaggle.com/competitions/google-smartphone-decimeter-challenge/discussion/262364\" target=\"_blank\">last year’s 10th place</a> <a href=\"https://www.kaggle.com/sai11fkaneko\" target=\"_blank\">@sai11fkaneko</a> used conv nets, and LGBMs  </p>\n<p><strong>7 - Hyperparameter Tune</strong><br>\nTuning your post processing methods and models can help you squeeze the best results out of the data. You can use the train baseline and ground truths to help you tune these models. In the notebook linked under the outlier correction and post-processing tabs, you can see how I used Bayesian Optimization to improve the results of the previous methods.<br>\n<a href=\"https://scikit-optimize.github.io/stable/\" target=\"_blank\">Documentation to skopt</a> <br>\n<a href=\"https://optuna.readthedocs.io/en/stable/\" target=\"_blank\">Documentation to optuna</a></p>\n<p><strong>8 - Validation with Train Files</strong><br>\nEverything you apply to your test set, you should also apply to your train set. Rather than choosing your post processing and parameters based on the leaderboard, I recommend evaluating with the train set. You can use the groundtruth.csv files to help you.<br>\nLast year I remember my private leaderboard rank took a very large drop because I overfit to the leaderboard. This year I’m going to try to avoid that.</p>\n<p>If you have any more tips or questions feel free to comment them.</p>",
  "messages": [
    {
      "id": 1780061,
      "postDate": "2022-05-07T04:09:12.030Z",
      "content": "<p><strong>1 - Understand the format of this dataset, objective, and evaluation</strong> <br>\nI recommend getting familiar with the structure of the data and going through the <a href=\"https://www.kaggle.com/competitions/smartphone-decimeter-2022/data\" target=\"_blank\">competition data tab</a> </p>\n<p>The goal of this competition is to predict the exact latitude and longitude coordinates of a car based on GNSS data coming from cell phones. The submission file will then be scored using the mean of the 50th and 95th percentile distance errors (<a href=\"https://www.kaggle.com/competitions/smartphone-decimeter-2022/overview/evaluation\" target=\"_blank\">see more details</a>)</p>\n<p><strong>2 - Make a baseline</strong><br>\nSince the baseline for this year's competition is in ECEF (Earth-Centered Earth-Fixed) coordinate system, the coordinate system must be converted to BLH to get the latitude and longitude submission baseline.<br>\n<a href=\"https://www.kaggle.com/saitodevel01\" target=\"_blank\">@saitodevel01</a> has published <a href=\"https://www.kaggle.com/code/saitodevel01/gsdc2-baseline-submission\" target=\"_blank\">an excellent notebook</a> demonstrating how to get to this baseline. <br>\nThe baseline this year seems to be a public score of 4.87</p>\n<p>After you have this baseline, you basically have the path the car took, but the signal is noisy. Now you must correct the baseline and find the true path.</p>\n<p><strong>3 - Correct Outliers</strong><br>\nSince we are working with data in time series paths, correcting outliers can be especially useful and should probably be one of your first steps. <a href=\"https://www.kaggle.com/code/ravishah1/gsdc2-savgol-filter-outlier-removal-with-bo\" target=\"_blank\">In my notebook</a>, I demonstrate a popular method of outlier correction I saw used in last year's competition.</p>\n<p><strong>4 - Smooth Baseline</strong><br>\nData smoothing uses an algorithm to remove noise from a data set. This can be very useful when working with gnss data. <a href=\"https://www.kaggle.com/code/ravishah1/gsdc2-savgol-filter-outlier-removal-with-bo\" target=\"_blank\">In my notebook</a>, I also demonstrate how to use the savgol filter to smooth the baseline. </p>\n<p>There are also many other forms of data smoothing you can try such as the Kalman filter which was popular in <a href=\"https://www.kaggle.com/code/emaerthin/demonstration-of-the-kalman-filter\" target=\"_blank\">public notebooks</a> such as this one by <a href=\"https://www.kaggle.com/emaerthin\" target=\"_blank\">@emaerthin</a> in last year's competition. </p>\n<p><strong>5 - Post Processing</strong><br>\nThere are several other types of post processing that you can use besides data smoothing. A lot of ideas such as snap to grid were used last year. Reading up on post processing for geospatial data may help give you an edge.<br>\nTo learn more about gnss data check out the resources in <a href=\"https://www.kaggle.com/chris62\" target=\"_blank\">@chris62</a> <a href=\"https://www.kaggle.com/competitions/smartphone-decimeter-2022/discussion/323057\" target=\"_blank\">discussion</a></p>\n<p><strong>6 - Machine Learning</strong><br>\nMachine learning may be a more challenging approach in this competition and post processing may be a simpler and more rewarding approach to start with. Nonetheless, using machine learning effectively may give you an edge in this competition. <br>\nSee how <a href=\"https://www.kaggle.com/competitions/google-smartphone-decimeter-challenge/discussion/262364\" target=\"_blank\">last year’s 10th place</a> <a href=\"https://www.kaggle.com/sai11fkaneko\" target=\"_blank\">@sai11fkaneko</a> used conv nets, and LGBMs  </p>\n<p><strong>7 - Hyperparameter Tune</strong><br>\nTuning your post processing methods and models can help you squeeze the best results out of the data. You can use the train baseline and ground truths to help you tune these models. In the notebook linked under the outlier correction and post-processing tabs, you can see how I used Bayesian Optimization to improve the results of the previous methods.<br>\n<a href=\"https://scikit-optimize.github.io/stable/\" target=\"_blank\">Documentation to skopt</a> <br>\n<a href=\"https://optuna.readthedocs.io/en/stable/\" target=\"_blank\">Documentation to optuna</a></p>\n<p><strong>8 - Validation with Train Files</strong><br>\nEverything you apply to your test set, you should also apply to your train set. Rather than choosing your post processing and parameters based on the leaderboard, I recommend evaluating with the train set. You can use the groundtruth.csv files to help you.<br>\nLast year I remember my private leaderboard rank took a very large drop because I overfit to the leaderboard. This year I’m going to try to avoid that.</p>\n<p>If you have any more tips or questions feel free to comment them.</p>",
      "rawMarkdown": "**1 - Understand the format of this dataset, objective, and evaluation** \nI recommend getting familiar with the structure of the data and going through the [competition data tab](https://www.kaggle.com/competitions/smartphone-decimeter-2022/data) \n\nThe goal of this competition is to predict the exact latitude and longitude coordinates of a car based on GNSS data coming from cell phones. The submission file will then be scored using the mean of the 50th and 95th percentile distance errors ([see more details](https://www.kaggle.com/competitions/smartphone-decimeter-2022/overview/evaluation))\n\n**2 - Make a baseline**\nSince the baseline for this year's competition is in ECEF (Earth-Centered Earth-Fixed) coordinate system, the coordinate system must be converted to BLH to get the latitude and longitude submission baseline.\n@saitodevel01 has published [an excellent notebook](https://www.kaggle.com/code/saitodevel01/gsdc2-baseline-submission) demonstrating how to get to this baseline. \nThe baseline this year seems to be a public score of 4.87\n\nAfter you have this baseline, you basically have the path the car took, but the signal is noisy. Now you must correct the baseline and find the true path.\n\n**3 - Correct Outliers**\nSince we are working with data in time series paths, correcting outliers can be especially useful and should probably be one of your first steps. [In my notebook](https://www.kaggle.com/code/ravishah1/gsdc2-savgol-filter-outlier-removal-with-bo), I demonstrate a popular method of outlier correction I saw used in last year's competition.\n\n**4 - Smooth Baseline**\nData smoothing uses an algorithm to remove noise from a data set. This can be very useful when working with gnss data. [In my notebook](https://www.kaggle.com/code/ravishah1/gsdc2-savgol-filter-outlier-removal-with-bo), I also demonstrate how to use the savgol filter to smooth the baseline. \n\nThere are also many other forms of data smoothing you can try such as the Kalman filter which was popular in [public notebooks](https://www.kaggle.com/code/emaerthin/demonstration-of-the-kalman-filter) such as this one by @emaerthin in last year's competition. \n\n**5 - Post Processing**\nThere are several other types of post processing that you can use besides data smoothing. A lot of ideas such as snap to grid were used last year. Reading up on post processing for geospatial data may help give you an edge.\nTo learn more about gnss data check out the resources in @chris62 [discussion](https://www.kaggle.com/competitions/smartphone-decimeter-2022/discussion/323057)\n\n**6 - Machine Learning**\nMachine learning may be a more challenging approach in this competition and post processing may be a simpler and more rewarding approach to start with. Nonetheless, using machine learning effectively may give you an edge in this competition. \nSee how [last year’s 10th place](https://www.kaggle.com/competitions/google-smartphone-decimeter-challenge/discussion/262364) @sai11fkaneko used conv nets, and LGBMs  \n\n**7 - Hyperparameter Tune**\nTuning your post processing methods and models can help you squeeze the best results out of the data. You can use the train baseline and ground truths to help you tune these models. In the notebook linked under the outlier correction and post-processing tabs, you can see how I used Bayesian Optimization to improve the results of the previous methods.\n[Documentation to skopt](https://scikit-optimize.github.io/stable/ ) \n[Documentation to optuna](https://optuna.readthedocs.io/en/stable/)\n\n**8 - Validation with Train Files**\nEverything you apply to your test set, you should also apply to your train set. Rather than choosing your post processing and parameters based on the leaderboard, I recommend evaluating with the train set. You can use the groundtruth.csv files to help you.\nLast year I remember my private leaderboard rank took a very large drop because I overfit to the leaderboard. This year I’m going to try to avoid that.\n\nIf you have any more tips or questions feel free to comment them.\n",
      "votes": 65
    },
    {
      "id": 1832768,
      "postDate": "2022-06-25T11:32:41.530Z",
      "content": "<p>Thanks for putting this together <a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">@ravishah1</a> , super useful stuff!</p>",
      "rawMarkdown": "Thanks for putting this together @ravishah1 , super useful stuff!",
      "votes": 1
    },
    {
      "id": 1817170,
      "postDate": "2022-06-11T02:18:13.917Z",
      "content": "<p>Thank you very much for the beautiful guidelines.<br>\nIt's very helpful.</p>",
      "rawMarkdown": "Thank you very much for the beautiful guidelines.\nIt's very helpful.",
      "votes": 1
    },
    {
      "id": 1797586,
      "postDate": "2022-05-22T05:41:24.877Z",
      "content": "<p>This is a really helpful guideline to approach the competition! I felt lost when I saw the dataset, but this source helped me understand the competition better! Thank you</p>",
      "rawMarkdown": "This is a really helpful guideline to approach the competition! I felt lost when I saw the dataset, but this source helped me understand the competition better! Thank you",
      "votes": 1
    },
    {
      "id": 1797372,
      "postDate": "2022-05-21T23:18:43.517Z",
      "content": "<p>It's an excellent beginner source, especially for someone to get into a quickstart in the competition. Thanks for this extremely curated post, even the links are great resources to learn 👍👍👍👍</p>",
      "rawMarkdown": "It's an excellent beginner source, especially for someone to get into a quickstart in the competition. Thanks for this extremely curated post, even the links are great resources to learn 👍👍👍👍",
      "votes": 1
    },
    {
      "id": 1795258,
      "postDate": "2022-05-19T15:23:30.617Z",
      "content": "<p>good roadmap for novice!</p>",
      "rawMarkdown": "good roadmap for novice!",
      "votes": 1
    },
    {
      "id": 1785357,
      "postDate": "2022-05-12T03:48:36.313Z",
      "content": "<p>Great explanation man!</p>",
      "rawMarkdown": "Great explanation man!",
      "votes": 1
    },
    {
      "id": 1850157,
      "postDate": "2022-07-10T06:51:04.887Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1834643,
      "postDate": "2022-06-27T05:30:02.780Z",
      "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">@ravishah1</a> !</p>",
      "rawMarkdown": "Thank you very much @ravishah1 !",
      "votes": 1
    },
    {
      "id": 1830603,
      "postDate": "2022-06-23T14:32:43.703Z",
      "content": "<p>Thank you! This is very helpful!</p>",
      "rawMarkdown": "Thank you! This is very helpful!",
      "votes": 1
    },
    {
      "id": 1824424,
      "postDate": "2022-06-18T09:56:38.870Z",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">@ravishah1</a> </p>",
      "rawMarkdown": "Thank you @ravishah1 ",
      "votes": 1
    },
    {
      "id": 1824176,
      "postDate": "2022-06-18T03:59:30.107Z",
      "content": "<p>Very helpful, thank you!</p>",
      "rawMarkdown": "Very helpful, thank you!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1832768,
      "author_name": "Pardeep Singh",
      "author_url": "",
      "post_date": "2022-06-25T11:32:41.530000",
      "content": "<p>Thanks for putting this together <a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">@ravishah1</a> , super useful stuff!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1817170,
      "author_name": "min fuka",
      "author_url": "",
      "post_date": "2022-06-11T02:18:13.917000",
      "content": "<p>Thank you very much for the beautiful guidelines.<br>\nIt's very helpful.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1797586,
      "author_name": "SOYOUNG PARK",
      "author_url": "",
      "post_date": "2022-05-22T05:41:24.877000",
      "content": "<p>This is a really helpful guideline to approach the competition! I felt lost when I saw the dataset, but this source helped me understand the competition better! Thank you</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1797372,
      "author_name": "Seemran Mishra",
      "author_url": "",
      "post_date": "2022-05-21T23:18:43.517000",
      "content": "<p>It's an excellent beginner source, especially for someone to get into a quickstart in the competition. Thanks for this extremely curated post, even the links are great resources to learn 👍👍👍👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1795258,
      "author_name": "MessiC",
      "author_url": "",
      "post_date": "2022-05-19T15:23:30.617000",
      "content": "<p>good roadmap for novice!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1785357,
      "author_name": "Bhavesh Bisht",
      "author_url": "",
      "post_date": "2022-05-12T03:48:36.313000",
      "content": "<p>Great explanation man!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1850157,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-10T06:51:04.887000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1834643,
      "author_name": "Abhishek Tiwari",
      "author_url": "",
      "post_date": "2022-06-27T05:30:02.780000",
      "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">@ravishah1</a> !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1830603,
      "author_name": "Chad Loh",
      "author_url": "",
      "post_date": "2022-06-23T14:32:43.703000",
      "content": "<p>Thank you! This is very helpful!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1824424,
      "author_name": "Pankaj Kumar",
      "author_url": "",
      "post_date": "2022-06-18T09:56:38.870000",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">@ravishah1</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1824176,
      "author_name": "Jakob",
      "author_url": "",
      "post_date": "2022-06-18T03:59:30.107000",
      "content": "<p>Very helpful, thank you!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1780061": "**1 - Understand the format of this dataset, objective, and evaluation** \nI recommend getting familiar with the structure of the data and going through the [competition data tab](https://www.kaggle.com/competitions/smartphone-decimeter-2022/data) \n\nThe goal of this competition is to predict the exact latitude and longitude coordinates of a car based on GNSS data coming from cell phones. The submission file will then be scored using the mean of the 50th and 95th percentile distance errors ([see more details](https://www.kaggle.com/competitions/smartphone-decimeter-2022/overview/evaluation))\n\n**2 - Make a baseline**\nSince the baseline for this year's competition is in ECEF (Earth-Centered Earth-Fixed) coordinate system, the coordinate system must be converted to BLH to get the latitude and longitude submission baseline.\n@saitodevel01 has published [an excellent notebook](https://www.kaggle.com/code/saitodevel01/gsdc2-baseline-submission) demonstrating how to get to this baseline. \nThe baseline this year seems to be a public score of 4.87\n\nAfter you have this baseline, you basically have the path the car took, but the signal is noisy. Now you must correct the baseline and find the true path.\n\n**3 - Correct Outliers**\nSince we are working with data in time series paths, correcting outliers can be especially useful and should probably be one of your first steps. [In my notebook](https://www.kaggle.com/code/ravishah1/gsdc2-savgol-filter-outlier-removal-with-bo), I demonstrate a popular method of outlier correction I saw used in last year's competition.\n\n**4 - Smooth Baseline**\nData smoothing uses an algorithm to remove noise from a data set. This can be very useful when working with gnss data. [In my notebook](https://www.kaggle.com/code/ravishah1/gsdc2-savgol-filter-outlier-removal-with-bo), I also demonstrate how to use the savgol filter to smooth the baseline. \n\nThere are also many other forms of data smoothing you can try such as the Kalman filter which was popular in [public notebooks](https://www.kaggle.com/code/emaerthin/demonstration-of-the-kalman-filter) such as this one by @emaerthin in last year's competition. \n\n**5 - Post Processing**\nThere are several other types of post processing that you can use besides data smoothing. A lot of ideas such as snap to grid were used last year. Reading up on post processing for geospatial data may help give you an edge.\nTo learn more about gnss data check out the resources in @chris62 [discussion](https://www.kaggle.com/competitions/smartphone-decimeter-2022/discussion/323057)\n\n**6 - Machine Learning**\nMachine learning may be a more challenging approach in this competition and post processing may be a simpler and more rewarding approach to start with. Nonetheless, using machine learning effectively may give you an edge in this competition. \nSee how [last year’s 10th place](https://www.kaggle.com/competitions/google-smartphone-decimeter-challenge/discussion/262364) @sai11fkaneko used conv nets, and LGBMs  \n\n**7 - Hyperparameter Tune**\nTuning your post processing methods and models can help you squeeze the best results out of the data. You can use the train baseline and ground truths to help you tune these models. In the notebook linked under the outlier correction and post-processing tabs, you can see how I used Bayesian Optimization to improve the results of the previous methods.\n[Documentation to skopt](https://scikit-optimize.github.io/stable/ ) \n[Documentation to optuna](https://optuna.readthedocs.io/en/stable/)\n\n**8 - Validation with Train Files**\nEverything you apply to your test set, you should also apply to your train set. Rather than choosing your post processing and parameters based on the leaderboard, I recommend evaluating with the train set. You can use the groundtruth.csv files to help you.\nLast year I remember my private leaderboard rank took a very large drop because I overfit to the leaderboard. This year I’m going to try to avoid that.\n\nIf you have any more tips or questions feel free to comment them.\n",
    "1832768": "Thanks for putting this together @ravishah1 , super useful stuff!",
    "1817170": "Thank you very much for the beautiful guidelines.\nIt's very helpful.",
    "1797586": "This is a really helpful guideline to approach the competition! I felt lost when I saw the dataset, but this source helped me understand the competition better! Thank you",
    "1797372": "It's an excellent beginner source, especially for someone to get into a quickstart in the competition. Thanks for this extremely curated post, even the links are great resources to learn 👍👍👍👍",
    "1795258": "good roadmap for novice!",
    "1785357": "Great explanation man!",
    "1850157": "",
    "1834643": "Thank you very much @ravishah1 !",
    "1830603": "Thank you! This is very helpful!",
    "1824424": "Thank you @ravishah1 ",
    "1824176": "Very helpful, thank you!"
  }
}