{
  "id": 261732,
  "title": "7th place solution",
  "url": "/competitions/google-smartphone-decimeter-challenge/discussion/261732",
  "author_name": "T88",
  "post_date": "2021-08-05T00:34:13.806000",
  "votes": 41,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First of all, I would like to thank host for organizing this competition. I participated in the competition as a soloist from start to finish, and although it was a very tough competition, but it was very meaningful as it gave me a chance to experience the interesting technology of GNSS.  </p>\n<p>Here is my solution.  </p>\n<h1>1.Baseline improving</h1>\n<p>I have rebuilt the baseline based on <a href=\"https://www.kaggle.com/hyperc/gsdc-reproducing-baseline-wls-on-one-measurement\" target=\"_blank\">this notebook</a>.  <br>\nFor isrbm, I used the median value for each phone-sat.  </p>\n<h2>1-1.Selecting a satellite</h2>\n<p>I improved the baseline by excluding satellites, which are a source of error, from the least-squares calculation.  </p>\n<p>For the filter condition, I mainly used the elevation angle. Signals from satellites with low elevation angles are excluded because they are strongly affected by various errors.  </p>\n<h2>1-2.Carrier smoothing</h2>\n<p>Pseudorange smoothing with Acumulated Delta Range (ADR), as described in <a href=\"https://www.kaggle.com/gymf123/onepager-tip-acumulated-delta-range-adr\" target=\"_blank\">this notebook</a>. The original pseudorange, and the previous pseudorange + ADR Diff mixed in a certain ratio to form the final pseudorange. ADR is relative but accurate, so it can be combined with the absolute value of pseudorange to improve the accuracy. Since ADR can have an accumulated value of zero due to cycle slip, I applied carrier smoothing only when AccumulatedDeltaRangeState = 25.  </p>\n<h1>2.Estimation of relative position</h1>\n<h2>2-1. Vehicle speed calculation using doppler shift</h2>\n<p>The relative velocity between the satellite and the vehicle can be determined by the frequency change of the signal (doppler shift). To determine the speed of a vehicle, we need the position of the satellite, the position of the vehicle, the speed of the satellite, the distance between the satellite and the vehicle, and the doppler shift. The position and speed of the satellite are given as data. The position of the vehicle and the distance between the satellite and the vehicle are obtained from the calculation results of Baseline improving (Need to subtract clkbias from pseudorange). The doppler shift is given as PseudorangeRateMetersPerSecond. The vehicle speed is then calculated by the least squares method using information from multiple satellites.  </p>\n<p>The vehicle speed (and the relative position calculated from it) obtained by this method was very accurate, and was a major factor in improving the score.  </p>\n<h2>2-2. ML prediction (add IMU data)</h2>\n<p>Since the relative positions obtained in 2-1 are missing in some places, I also combined IMU sensor data to create a machine learning model to supplement them. I built a prediction model in lightGBM with lag and rolling features of the IMU and vehicle velocity.</p>\n<h1>3. Reject outlier</h1>\n<p>There are some outliers in the baseline, which I will remove.</p>\n<h2>3-1. Abnormally high speeds</h2>\n<p>Exclude points that have a very large distance from the previous and next point. </p>\n<h2>3-2. Based on ground truth</h2>\n<p>Since some areas have overlapping test and train paths, I were able to use the ground truth of the train to determine the outlier. The closest distance to the ground truth data was calculated for each point, and those above the threshold were removed as outlier.  </p>\n<h2>3-3. Based on reference point calculated by relative position</h2>\n<p>Since there are many test data that have paths that do not exist in train, the 3-2 method can only be used in a very limited way. To solve this problem, I created a reference point that can be used as an alternative to ground truth.  <br>\nFor this, I used the relative positions calculated in 2. Starting from the coordinate point at each time, the coordinates before and after a certain time are calculated based on the accumulated relative values. By sliding this process at each point in time, a large number of estimates can be obtained at each time. The accuracy of these estimates is highly dependent on the accuracy of the absolute coordinates of the starting point. If the starting point is an outlier, the estimated value will also be an outlier, but this is not frequent and the effect can be eliminated by clipping the estimated value. We then calculated a threshold value from the mean and standard deviation of the estimated values at each point, and used it to remove the outlier values.  </p>\n<h1>4. Post process</h1>\n<h2>4-1. kalman smoothing</h2>\n<p>I used <a href=\"https://www.kaggle.com/emaerthin/demonstration-of-the-kalman-filter\" target=\"_blank\">this notebook</a> as is.  </p>\n<h2>4-2. Processing the speed0 period</h2>\n<p>As discussed in <a href=\"https://www.kaggle.com/t88take/gsdc-eda-error-when-stopping\" target=\"_blank\">this notebook</a>, there is a tendency for absolute coordinates to be highly scattered when the car is stopping. To solve this problem, I created a model to predict stops, and replaced the continuous periods predicted as stops with the average of those data.  </p>\n<h2>4-3. Cost minimization</h2>\n<p>I used the <a href=\"https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization\" target=\"_blank\">cost_minimization notebook</a>. Improved the accuracy of the absolute coordinates based on the relative position obtained in 2.  </p>\n<h2>4-4. Position shift</h2>\n<p>I used <a href=\"https://www.kaggle.com/wrrosa/gsdc-position-shift\" target=\"_blank\">this notebook</a> as is.</p>\n<h2>4-5. Weighted phones mean</h2>\n<p>As described in <a href=\"https://www.kaggle.com/t88take/gsdc-phones-mean-prediction\" target=\"_blank\">this notebook</a>, this process averages the values of multiple phones in the same collection. It has been improved from the published version and changed to add weight to each phone model.  </p>\n<h1>Area grouping</h1>\n<p>I added a little logic to the kNN introduced <a href=\"https://www.kaggle.com/columbia2131/area-knn-prediction-train-hand-label\" target=\"_blank\">here</a>, and implemented grouping based on the degree of path matching with the train. Each collection was divided into five groups, and the hyperparameters and order of processing were adjusted for each group.  </p>",
  "messages": [
    {
      "id": 1449757,
      "postDate": "2021-08-05T00:34:13.807Z",
      "content": "<p>First of all, I would like to thank host for organizing this competition. I participated in the competition as a soloist from start to finish, and although it was a very tough competition, but it was very meaningful as it gave me a chance to experience the interesting technology of GNSS.  </p>\n<p>Here is my solution.  </p>\n<h1>1.Baseline improving</h1>\n<p>I have rebuilt the baseline based on <a href=\"https://www.kaggle.com/hyperc/gsdc-reproducing-baseline-wls-on-one-measurement\" target=\"_blank\">this notebook</a>.  <br>\nFor isrbm, I used the median value for each phone-sat.  </p>\n<h2>1-1.Selecting a satellite</h2>\n<p>I improved the baseline by excluding satellites, which are a source of error, from the least-squares calculation.  </p>\n<p>For the filter condition, I mainly used the elevation angle. Signals from satellites with low elevation angles are excluded because they are strongly affected by various errors.  </p>\n<h2>1-2.Carrier smoothing</h2>\n<p>Pseudorange smoothing with Acumulated Delta Range (ADR), as described in <a href=\"https://www.kaggle.com/gymf123/onepager-tip-acumulated-delta-range-adr\" target=\"_blank\">this notebook</a>. The original pseudorange, and the previous pseudorange + ADR Diff mixed in a certain ratio to form the final pseudorange. ADR is relative but accurate, so it can be combined with the absolute value of pseudorange to improve the accuracy. Since ADR can have an accumulated value of zero due to cycle slip, I applied carrier smoothing only when AccumulatedDeltaRangeState = 25.  </p>\n<h1>2.Estimation of relative position</h1>\n<h2>2-1. Vehicle speed calculation using doppler shift</h2>\n<p>The relative velocity between the satellite and the vehicle can be determined by the frequency change of the signal (doppler shift). To determine the speed of a vehicle, we need the position of the satellite, the position of the vehicle, the speed of the satellite, the distance between the satellite and the vehicle, and the doppler shift. The position and speed of the satellite are given as data. The position of the vehicle and the distance between the satellite and the vehicle are obtained from the calculation results of Baseline improving (Need to subtract clkbias from pseudorange). The doppler shift is given as PseudorangeRateMetersPerSecond. The vehicle speed is then calculated by the least squares method using information from multiple satellites.  </p>\n<p>The vehicle speed (and the relative position calculated from it) obtained by this method was very accurate, and was a major factor in improving the score.  </p>\n<h2>2-2. ML prediction (add IMU data)</h2>\n<p>Since the relative positions obtained in 2-1 are missing in some places, I also combined IMU sensor data to create a machine learning model to supplement them. I built a prediction model in lightGBM with lag and rolling features of the IMU and vehicle velocity.</p>\n<h1>3. Reject outlier</h1>\n<p>There are some outliers in the baseline, which I will remove.</p>\n<h2>3-1. Abnormally high speeds</h2>\n<p>Exclude points that have a very large distance from the previous and next point. </p>\n<h2>3-2. Based on ground truth</h2>\n<p>Since some areas have overlapping test and train paths, I were able to use the ground truth of the train to determine the outlier. The closest distance to the ground truth data was calculated for each point, and those above the threshold were removed as outlier.  </p>\n<h2>3-3. Based on reference point calculated by relative position</h2>\n<p>Since there are many test data that have paths that do not exist in train, the 3-2 method can only be used in a very limited way. To solve this problem, I created a reference point that can be used as an alternative to ground truth.  <br>\nFor this, I used the relative positions calculated in 2. Starting from the coordinate point at each time, the coordinates before and after a certain time are calculated based on the accumulated relative values. By sliding this process at each point in time, a large number of estimates can be obtained at each time. The accuracy of these estimates is highly dependent on the accuracy of the absolute coordinates of the starting point. If the starting point is an outlier, the estimated value will also be an outlier, but this is not frequent and the effect can be eliminated by clipping the estimated value. We then calculated a threshold value from the mean and standard deviation of the estimated values at each point, and used it to remove the outlier values.  </p>\n<h1>4. Post process</h1>\n<h2>4-1. kalman smoothing</h2>\n<p>I used <a href=\"https://www.kaggle.com/emaerthin/demonstration-of-the-kalman-filter\" target=\"_blank\">this notebook</a> as is.  </p>\n<h2>4-2. Processing the speed0 period</h2>\n<p>As discussed in <a href=\"https://www.kaggle.com/t88take/gsdc-eda-error-when-stopping\" target=\"_blank\">this notebook</a>, there is a tendency for absolute coordinates to be highly scattered when the car is stopping. To solve this problem, I created a model to predict stops, and replaced the continuous periods predicted as stops with the average of those data.  </p>\n<h2>4-3. Cost minimization</h2>\n<p>I used the <a href=\"https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization\" target=\"_blank\">cost_minimization notebook</a>. Improved the accuracy of the absolute coordinates based on the relative position obtained in 2.  </p>\n<h2>4-4. Position shift</h2>\n<p>I used <a href=\"https://www.kaggle.com/wrrosa/gsdc-position-shift\" target=\"_blank\">this notebook</a> as is.</p>\n<h2>4-5. Weighted phones mean</h2>\n<p>As described in <a href=\"https://www.kaggle.com/t88take/gsdc-phones-mean-prediction\" target=\"_blank\">this notebook</a>, this process averages the values of multiple phones in the same collection. It has been improved from the published version and changed to add weight to each phone model.  </p>\n<h1>Area grouping</h1>\n<p>I added a little logic to the kNN introduced <a href=\"https://www.kaggle.com/columbia2131/area-knn-prediction-train-hand-label\" target=\"_blank\">here</a>, and implemented grouping based on the degree of path matching with the train. Each collection was divided into five groups, and the hyperparameters and order of processing were adjusted for each group.  </p>",
      "rawMarkdown": "First of all, I would like to thank host for organizing this competition. I participated in the competition as a soloist from start to finish, and although it was a very tough competition, but it was very meaningful as it gave me a chance to experience the interesting technology of GNSS.  \n\nHere is my solution.  \n\n# 1.Baseline improving\nI have rebuilt the baseline based on [this notebook](https://www.kaggle.com/hyperc/gsdc-reproducing-baseline-wls-on-one-measurement).  \nFor isrbm, I used the median value for each phone-sat.  \n\n## 1-1.Selecting a satellite\nI improved the baseline by excluding satellites, which are a source of error, from the least-squares calculation.  \n\nFor the filter condition, I mainly used the elevation angle. Signals from satellites with low elevation angles are excluded because they are strongly affected by various errors.  \n\n## 1-2.Carrier smoothing\nPseudorange smoothing with Acumulated Delta Range (ADR), as described in [this notebook](https://www.kaggle.com/gymf123/onepager-tip-acumulated-delta-range-adr). The original pseudorange, and the previous pseudorange + ADR Diff mixed in a certain ratio to form the final pseudorange. ADR is relative but accurate, so it can be combined with the absolute value of pseudorange to improve the accuracy. Since ADR can have an accumulated value of zero due to cycle slip, I applied carrier smoothing only when AccumulatedDeltaRangeState = 25.  \n\n# 2.Estimation of relative position\n## 2-1. Vehicle speed calculation using doppler shift\nThe relative velocity between the satellite and the vehicle can be determined by the frequency change of the signal (doppler shift). To determine the speed of a vehicle, we need the position of the satellite, the position of the vehicle, the speed of the satellite, the distance between the satellite and the vehicle, and the doppler shift. The position and speed of the satellite are given as data. The position of the vehicle and the distance between the satellite and the vehicle are obtained from the calculation results of Baseline improving (Need to subtract clkbias from pseudorange). The doppler shift is given as PseudorangeRateMetersPerSecond. The vehicle speed is then calculated by the least squares method using information from multiple satellites.  \n\nThe vehicle speed (and the relative position calculated from it) obtained by this method was very accurate, and was a major factor in improving the score.  \n\n## 2-2. ML prediction (add IMU data) \nSince the relative positions obtained in 2-1 are missing in some places, I also combined IMU sensor data to create a machine learning model to supplement them. I built a prediction model in lightGBM with lag and rolling features of the IMU and vehicle velocity.\n\n# 3. Reject outlier\nThere are some outliers in the baseline, which I will remove.\n\n## 3-1. Abnormally high speeds  \nExclude points that have a very large distance from the previous and next point. \n\n## 3-2. Based on ground truth\nSince some areas have overlapping test and train paths, I were able to use the ground truth of the train to determine the outlier. The closest distance to the ground truth data was calculated for each point, and those above the threshold were removed as outlier.  \n\n## 3-3. Based on reference point calculated by relative position \nSince there are many test data that have paths that do not exist in train, the 3-2 method can only be used in a very limited way. To solve this problem, I created a reference point that can be used as an alternative to ground truth.  \nFor this, I used the relative positions calculated in 2. Starting from the coordinate point at each time, the coordinates before and after a certain time are calculated based on the accumulated relative values. By sliding this process at each point in time, a large number of estimates can be obtained at each time. The accuracy of these estimates is highly dependent on the accuracy of the absolute coordinates of the starting point. If the starting point is an outlier, the estimated value will also be an outlier, but this is not frequent and the effect can be eliminated by clipping the estimated value. We then calculated a threshold value from the mean and standard deviation of the estimated values at each point, and used it to remove the outlier values.  \n\n# 4. Post process\n## 4-1. kalman smoothing\nI used [this notebook](https://www.kaggle.com/emaerthin/demonstration-of-the-kalman-filter) as is.  \n\n## 4-2. Processing the speed0 period\nAs discussed in [this notebook](https://www.kaggle.com/t88take/gsdc-eda-error-when-stopping), there is a tendency for absolute coordinates to be highly scattered when the car is stopping. To solve this problem, I created a model to predict stops, and replaced the continuous periods predicted as stops with the average of those data.  \n\n## 4-3. Cost minimization\nI used the [cost_minimization notebook](https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization). Improved the accuracy of the absolute coordinates based on the relative position obtained in 2.  \n\n## 4-4. Position shift\nI used [this notebook](https://www.kaggle.com/wrrosa/gsdc-position-shift) as is.\n\n## 4-5. Weighted phones mean\nAs described in [this notebook](https://www.kaggle.com/t88take/gsdc-phones-mean-prediction), this process averages the values of multiple phones in the same collection. It has been improved from the published version and changed to add weight to each phone model.  \n\n# Area grouping\nI added a little logic to the kNN introduced [here](https://www.kaggle.com/columbia2131/area-knn-prediction-train-hand-label), and implemented grouping based on the degree of path matching with the train. Each collection was divided into five groups, and the hyperparameters and order of processing were adjusted for each group.  ",
      "votes": 41
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1449757": "First of all, I would like to thank host for organizing this competition. I participated in the competition as a soloist from start to finish, and although it was a very tough competition, but it was very meaningful as it gave me a chance to experience the interesting technology of GNSS.  \n\nHere is my solution.  \n\n# 1.Baseline improving\nI have rebuilt the baseline based on [this notebook](https://www.kaggle.com/hyperc/gsdc-reproducing-baseline-wls-on-one-measurement).  \nFor isrbm, I used the median value for each phone-sat.  \n\n## 1-1.Selecting a satellite\nI improved the baseline by excluding satellites, which are a source of error, from the least-squares calculation.  \n\nFor the filter condition, I mainly used the elevation angle. Signals from satellites with low elevation angles are excluded because they are strongly affected by various errors.  \n\n## 1-2.Carrier smoothing\nPseudorange smoothing with Acumulated Delta Range (ADR), as described in [this notebook](https://www.kaggle.com/gymf123/onepager-tip-acumulated-delta-range-adr). The original pseudorange, and the previous pseudorange + ADR Diff mixed in a certain ratio to form the final pseudorange. ADR is relative but accurate, so it can be combined with the absolute value of pseudorange to improve the accuracy. Since ADR can have an accumulated value of zero due to cycle slip, I applied carrier smoothing only when AccumulatedDeltaRangeState = 25.  \n\n# 2.Estimation of relative position\n## 2-1. Vehicle speed calculation using doppler shift\nThe relative velocity between the satellite and the vehicle can be determined by the frequency change of the signal (doppler shift). To determine the speed of a vehicle, we need the position of the satellite, the position of the vehicle, the speed of the satellite, the distance between the satellite and the vehicle, and the doppler shift. The position and speed of the satellite are given as data. The position of the vehicle and the distance between the satellite and the vehicle are obtained from the calculation results of Baseline improving (Need to subtract clkbias from pseudorange). The doppler shift is given as PseudorangeRateMetersPerSecond. The vehicle speed is then calculated by the least squares method using information from multiple satellites.  \n\nThe vehicle speed (and the relative position calculated from it) obtained by this method was very accurate, and was a major factor in improving the score.  \n\n## 2-2. ML prediction (add IMU data) \nSince the relative positions obtained in 2-1 are missing in some places, I also combined IMU sensor data to create a machine learning model to supplement them. I built a prediction model in lightGBM with lag and rolling features of the IMU and vehicle velocity.\n\n# 3. Reject outlier\nThere are some outliers in the baseline, which I will remove.\n\n## 3-1. Abnormally high speeds  \nExclude points that have a very large distance from the previous and next point. \n\n## 3-2. Based on ground truth\nSince some areas have overlapping test and train paths, I were able to use the ground truth of the train to determine the outlier. The closest distance to the ground truth data was calculated for each point, and those above the threshold were removed as outlier.  \n\n## 3-3. Based on reference point calculated by relative position \nSince there are many test data that have paths that do not exist in train, the 3-2 method can only be used in a very limited way. To solve this problem, I created a reference point that can be used as an alternative to ground truth.  \nFor this, I used the relative positions calculated in 2. Starting from the coordinate point at each time, the coordinates before and after a certain time are calculated based on the accumulated relative values. By sliding this process at each point in time, a large number of estimates can be obtained at each time. The accuracy of these estimates is highly dependent on the accuracy of the absolute coordinates of the starting point. If the starting point is an outlier, the estimated value will also be an outlier, but this is not frequent and the effect can be eliminated by clipping the estimated value. We then calculated a threshold value from the mean and standard deviation of the estimated values at each point, and used it to remove the outlier values.  \n\n# 4. Post process\n## 4-1. kalman smoothing\nI used [this notebook](https://www.kaggle.com/emaerthin/demonstration-of-the-kalman-filter) as is.  \n\n## 4-2. Processing the speed0 period\nAs discussed in [this notebook](https://www.kaggle.com/t88take/gsdc-eda-error-when-stopping), there is a tendency for absolute coordinates to be highly scattered when the car is stopping. To solve this problem, I created a model to predict stops, and replaced the continuous periods predicted as stops with the average of those data.  \n\n## 4-3. Cost minimization\nI used the [cost_minimization notebook](https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization). Improved the accuracy of the absolute coordinates based on the relative position obtained in 2.  \n\n## 4-4. Position shift\nI used [this notebook](https://www.kaggle.com/wrrosa/gsdc-position-shift) as is.\n\n## 4-5. Weighted phones mean\nAs described in [this notebook](https://www.kaggle.com/t88take/gsdc-phones-mean-prediction), this process averages the values of multiple phones in the same collection. It has been improved from the published version and changed to add weight to each phone model.  \n\n# Area grouping\nI added a little logic to the kNN introduced [here](https://www.kaggle.com/columbia2131/area-knn-prediction-train-hand-label), and implemented grouping based on the degree of path matching with the train. Each collection was divided into five groups, and the hyperparameters and order of processing were adjusted for each group.  "
  }
}