{
  "id": 340692,
  "title": "5th place solution",
  "url": "/competitions/smartphone-decimeter-2022/writeups/a-saito-5th-place-solution",
  "author_name": "",
  "post_date": "2022-07-30T14:07:57.210226100Z",
  "votes": 44,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Thank you to hosts for hosting this great competition. <br>\nAt the start of this competition, I was concerned that this competition would not attract many Kagglers, because machine learning was not very effective for GNSS data in the last year competition, but I am surprised that the competition was at very high level in the end. <br>\nI would like to publish my solution here:</p>\n<h1>Summary</h1>\n<p>I followed the <a href=\"https://www.kaggle.com/taroz1461\" target=\"_blank\">@taroz1461</a>'s <a href=\"https://www.kaggle.com/competitions/google-smartphone-decimeter-challenge/discussion/262406\" target=\"_blank\">1st place solution from last year competition</a> and combined it with <a href=\"https://www.kaggle.com/competitions/google-smartphone-decimeter-challenge/discussion/261959\" target=\"_blank\">my ideas of the 5th place solution</a>.</p>\n<p>The key point of the solution is to formulate and solve the integration of GNSS observations (pseudorange, pseudorange-rate(doppler), ADR) and trajectory smoothing as a single optimization problem. This provides following advantages: accuracy of satellite positioning due to satellite constellation is automatically taken into account in trajectory smoothing, and satellite positioning can be continued even if the number of satellites is temporarily less than required. In addition, by adding switchable constraints to all GNSS observations, satellite positioning becomes robust.</p>\n<p>Optimization was done by writing an evaluation function and a custom solver for handling equality constraints using TensorFlow (not NN). The flexibility of the deep learning framework was helpful for quickly testing ideas, but most ideas did not contribute to accuracy improvement in this competition. In the end, optimization part was fairly simple.</p>\n<p><img src=\"https://user-images.githubusercontent.com/309785/181914831-1e6e3068-4046-4468-9070-dabc8d5975ef.png\" alt=\"summary\"></p>\n<h1>Input Data &amp; Preprocessing</h1>\n<ul>\n<li><p>Raw GNSS log (supplemental/gnss_log.txt): <br>\ndevice_gnss.csv had null lines for SignalType, satellite position, etc. So I used gnss_log.txt and added columns to the obtained data-frame.</p></li>\n<li><p>Observation of pseudorange of base stations:<br>\nI used <a href=\"https://www.unavco.org/data/gps-gnss/gps-gnss.html\" target=\"_blank\">UNAVCO's</a> 15-second rate observation data and calculated the residual between the base station pseudorange and distance between satellite and base station. Although it is not necessary to correct for ionospheric and tropospheric delays in relative positioning, I corrected for ionospheric and tropospheric delays for both base stations and smartphone, since I also used base station data with long baseline length when performing multiple base station ensembles.</p></li>\n<li><p>Precise ephemeris:<br>\nI used satellite orbit data from <a href=\"https://cddis.nasa.gov/Data_and_Derived_Products/GNSS/gnss_mgex_products.html\" target=\"_blank\">GNSS IGS MGEX Products</a> to obtain the satellite's position and velocity. Although the accuracy of the satellite position is not so critical in relative positioning, it was easier to interpolate the position data from the precise ephemeris than to calculate the satellite orbit from the navigation message.</p></li>\n<li><p>Base station location data:<br>\nI downloaded time series data of base station location data from <a href=\"http://geodesy.unr.edu/index.php\" target=\"_blank\">Nevada Geodetic Laboratory</a>, and used the average over the test data period (2021/4/28 - 2022/4/25) as the base station location.</p></li>\n<li><p><a href=\"https://cddis.nasa.gov/Data_and_Derived_Products/GNSS/atmospheric_products.html#iono\" target=\"_blank\">Ionospheric delay data</a>:<br>\nIt was used to compensate for the difference of ionospheric delay between the base station location and the smartphone location. The contribution to the final accuracy of using external data on ionospheric delay should be small, since the ionospheric delay should be mostly cancelled by relative positioning.</p></li>\n<li><p>WlsPosition in device_gnss.csv:<br>\nI used the smoothed results of the baseline locations provided by the host to pre-compute the Sagnac effect, ionospheric delay, and tropospheric delay.</p></li>\n<li><p>I did not use IMU sensor data at all in this competition because sensor data is noisy and irregular. I did a lot of work on signal processing of sensor data in the last year competition, but it did not contribute much to the improvement of my score.</p></li>\n<li><p>Since I did not use machine learning, I used ground-truth data only to check accuracy of prediction results.</p></li>\n</ul>\n<h1>Optimization</h1>\n<h2>State-space representation of smartphone trajectory</h2>\n<p>The position, velocity, acceleration, and jerk (derivative of acceleration) of the smartphone are represented in the ECEF coordinate system, and discretized state equations are created assuming that the jerk is piecewise constant.</p>\n<p><img src=\"https://user-images.githubusercontent.com/309785/181914830-93f67930-4d52-4b62-a6b1-de88f09867cd.png\" alt=\"state_equation\"></p>\n<p>Under this assumption, the position of the smartphone is represented by a piecewise cubic polynomial, so the position of the smartphone at any time can be accurately calculated by third-order Hermite interpolation using the position and velocity of collocation points (same for velocity).</p>\n<p>Using the position and velocity of the smartphone calculated by this method and the position and velocity of the satellites, the distance between the satellites and the smartphone and its derivative are calculated, and the trajectory of the smartphone is calculated by minimizing the residual between them and the GNSS observation data.</p>\n<h2>Position smoothing via quadratic programing</h2>\n<p>Smoothing is performed on the positions obtained by the weighted least squares method of psedurorange by solving a quadratic programming problem with the sum of the squares of the position residuals and jerks as the evaluation function.<br>\nThe coefficients of the position observation matrix are calculated based on Hermite interpolation of the smartphone position based on the satellite positioning time.</p>\n<p>In last year's competition, I added velocity and acceleration terms to the evaluation function of a similar quadratic programming problem to improve accuracy, but in this competition, I did not use doppler and ADR observations at this point because the purpose of quadratic programming problems is to obtain initial values for the global optimization that follows.</p>\n<h2>Global optimization</h2>\n<p><img src=\"https://user-images.githubusercontent.com/309785/181914822-b29b3c62-4973-4320-b86c-c1ddf800c332.png\" alt=\"computation\"></p>\n<p>The evaluation function is calculated by the computation graph shown above by TensorFlow.<br>\nIn this figure, the unknown variables (trainable variables) of the optimization are shown in red, pre-computable constants in blue. The inter-signal range biases are parameters shared to all times for each SignalType.</p>\n<p>Due to the availability of base station observations, only GPS, GLONASS, and Galileo were used for pseudorange observation data. For the doppler and ADR observations, I used data from all constellations, including QZSS and BeiDou.<br>\nIn addition, the weighting of residuals by RecievedSvTimeUncertaintyNanos, PseudoRangeUncertaintyMetersPerSecond, and AcuumulatedDeltaRangeUncertaintyMeters contributed to accuracy improvement.</p>\n<p>To take into account the equality constraints of the state equation, augmented Lagrangian method is used to optimize the evaluation function defined by the above computation graph.</p>\n<p><img src=\"https://user-images.githubusercontent.com/309785/181914819-7604ffec-7d19-4940-a3f0-bcd261917516.png\" alt=\"augmented_Lagrangian\"></p>\n<p>Where f(x) is the evaluation function and h(x) is the equality constraint corresponding to the above state equations.<br>\nMore specifically, the following steps are repeated alternately: <br>\n(1) minimizing the augmented Lagrangian by Adam optimizer.<br>\n(2) updating the dual variables of the equality constraints (λ).</p>\n<p>Convergence of this optimization is slow! But, it can robustly compute solutions with good accuracy.</p>\n<p>Note that double precision floating point is used in all computations to eliminate floating point rounding errors. For this reason, optimization is performed using the CPU instead of the GPU.</p>\n<h1>Workaround for pseudorange-rate of XiaomiMi8</h1>\n<p><img src=\"https://user-images.githubusercontent.com/309785/181914836-97fa4def-3d2f-4d55-9534-33fd43fc8463.png\" alt=\"velocity_XiaomiMi8\"></p>\n<p>This figure shows comparison between ground-truth velocity and estimated velocity by weighted least-squares method using pseudorange-rate in 2021-12-07-US-LAX-1/XiaomiMi8.<br>\nThis shows that there is a time discrepancy of about 0.5 seconds between the estimated velocity and the ground-truth (WHY?). Therefore, correcting for this time discrepancy improves the accuracy of position estimation.<br>\nIn my solution, this problem is addressed by a correction term (acceleration × time delta) for smartphone velocity and adding the time delta parameter to the optimization variables. (This workaround is applied to data other than XiaomiMi8.) This was quite effective in improving public LB, because there are many XiaomiMi8 in public test data.</p>\n<p><img src=\"https://user-images.githubusercontent.com/309785/181914823-1272523e-2a04-4d4a-b6d8-e0bf64797501.png\" alt=\"modify_velocity\"></p>\n<h1>Ensemble of multiple base station solutions</h1>\n<p><img src=\"https://user-images.githubusercontent.com/309785/181917131-42d34aa8-8e9a-4fa3-b9ca-c06204c6fe58.jpg\" alt=\"google-earth\"></p>\n<p>I had originally planed to interpolate solutions from multiple base stations based on the location of base stations and smartphone, but I found that the correlation between baseline length and position error was not so clear, and simply averaging the solutions from multiple base stations was slightly more accurate. The reason for this is assumed to be that the pseudorange data of the base stations also contains noise of about 0.5 to 1.0 m. Therefore, the median of the solutions of multiple base stations was used as the final submission, after correcting the difference between the ionospheric and tropospheric delays at the base stations and smartphone locations using external data and the model.</p>",
  "messages": [
    {
      "id": "1877297",
      "postDate": "07/30/2022 14:07:57",
      "content": "<p>Thank you to hosts for hosting this great competition. <br>\nAt the start of this competition, I was concerned that this competition would not attract many Kagglers, because machine learning was not very effective for GNSS data in the last year competition, but I am surprised that the competition was at very high level in the end. <br>\nI would like to publish my solution here:</p>\n<h1>Summary</h1>\n<p>I followed the <a href=\"https://www.kaggle.com/taroz1461\" target=\"_blank\">@taroz1461</a>'s <a href=\"https://www.kaggle.com/competitions/google-smartphone-decimeter-challenge/discussion/262406\" target=\"_blank\">1st place solution from last year competition</a> and combined it with <a href=\"https://www.kaggle.com/competitions/google-smartphone-decimeter-challenge/discussion/261959\" target=\"_blank\">my ideas of the 5th place solution</a>.</p>\n<p>The key point of the solution is to formulate and solve the integration of GNSS observations (pseudorange, pseudorange-rate(doppler), ADR) and trajectory smoothing as a single optimization problem. This provides following advantages: accuracy of satellite positioning due to satellite constellation is automatically taken into account in trajectory smoothing, and satellite positioning can be continued even if the number of satellites is temporarily less than required. In addition, by adding switchable constraints to all GNSS observations, satellite positioning becomes robust.</p>\n<p>Optimization was done by writing an evaluation function and a custom solver for handling equality constraints using TensorFlow (not NN). The flexibility of the deep learning framework was helpful for quickly testing ideas, but most ideas did not contribute to accuracy improvement in this competition. In the end, optimization part was fairly simple.</p>\n<p><img src=\"https://user-images.githubusercontent.com/309785/181914831-1e6e3068-4046-4468-9070-dabc8d5975ef.png\" alt=\"summary\"></p>\n<h1>Input Data &amp; Preprocessing</h1>\n<ul>\n<li><p>Raw GNSS log (supplemental/gnss_log.txt): <br>\ndevice_gnss.csv had null lines for SignalType, satellite position, etc. So I used gnss_log.txt and added columns to the obtained data-frame.</p></li>\n<li><p>Observation of pseudorange of base stations:<br>\nI used <a href=\"https://www.unavco.org/data/gps-gnss/gps-gnss.html\" target=\"_blank\">UNAVCO's</a> 15-second rate observation data and calculated the residual between the base station pseudorange and distance between satellite and base station. Although it is not necessary to correct for ionospheric and tropospheric delays in relative positioning, I corrected for ionospheric and tropospheric delays for both base stations and smartphone, since I also used base station data with long baseline length when performing multiple base station ensembles.</p></li>\n<li><p>Precise ephemeris:<br>\nI used satellite orbit data from <a href=\"https://cddis.nasa.gov/Data_and_Derived_Products/GNSS/gnss_mgex_products.html\" target=\"_blank\">GNSS IGS MGEX Products</a> to obtain the satellite's position and velocity. Although the accuracy of the satellite position is not so critical in relative positioning, it was easier to interpolate the position data from the precise ephemeris than to calculate the satellite orbit from the navigation message.</p></li>\n<li><p>Base station location data:<br>\nI downloaded time series data of base station location data from <a href=\"http://geodesy.unr.edu/index.php\" target=\"_blank\">Nevada Geodetic Laboratory</a>, and used the average over the test data period (2021/4/28 - 2022/4/25) as the base station location.</p></li>\n<li><p><a href=\"https://cddis.nasa.gov/Data_and_Derived_Products/GNSS/atmospheric_products.html#iono\" target=\"_blank\">Ionospheric delay data</a>:<br>\nIt was used to compensate for the difference of ionospheric delay between the base station location and the smartphone location. The contribution to the final accuracy of using external data on ionospheric delay should be small, since the ionospheric delay should be mostly cancelled by relative positioning.</p></li>\n<li><p>WlsPosition in device_gnss.csv:<br>\nI used the smoothed results of the baseline locations provided by the host to pre-compute the Sagnac effect, ionospheric delay, and tropospheric delay.</p></li>\n<li><p>I did not use IMU sensor data at all in this competition because sensor data is noisy and irregular. I did a lot of work on signal processing of sensor data in the last year competition, but it did not contribute much to the improvement of my score.</p></li>\n<li><p>Since I did not use machine learning, I used ground-truth data only to check accuracy of prediction results.</p></li>\n</ul>\n<h1>Optimization</h1>\n<h2>State-space representation of smartphone trajectory</h2>\n<p>The position, velocity, acceleration, and jerk (derivative of acceleration) of the smartphone are represented in the ECEF coordinate system, and discretized state equations are created assuming that the jerk is piecewise constant.</p>\n<p><img src=\"https://user-images.githubusercontent.com/309785/181914830-93f67930-4d52-4b62-a6b1-de88f09867cd.png\" alt=\"state_equation\"></p>\n<p>Under this assumption, the position of the smartphone is represented by a piecewise cubic polynomial, so the position of the smartphone at any time can be accurately calculated by third-order Hermite interpolation using the position and velocity of collocation points (same for velocity).</p>\n<p>Using the position and velocity of the smartphone calculated by this method and the position and velocity of the satellites, the distance between the satellites and the smartphone and its derivative are calculated, and the trajectory of the smartphone is calculated by minimizing the residual between them and the GNSS observation data.</p>\n<h2>Position smoothing via quadratic programing</h2>\n<p>Smoothing is performed on the positions obtained by the weighted least squares method of psedurorange by solving a quadratic programming problem with the sum of the squares of the position residuals and jerks as the evaluation function.<br>\nThe coefficients of the position observation matrix are calculated based on Hermite interpolation of the smartphone position based on the satellite positioning time.</p>\n<p>In last year's competition, I added velocity and acceleration terms to the evaluation function of a similar quadratic programming problem to improve accuracy, but in this competition, I did not use doppler and ADR observations at this point because the purpose of quadratic programming problems is to obtain initial values for the global optimization that follows.</p>\n<h2>Global optimization</h2>\n<p><img src=\"https://user-images.githubusercontent.com/309785/181914822-b29b3c62-4973-4320-b86c-c1ddf800c332.png\" alt=\"computation\"></p>\n<p>The evaluation function is calculated by the computation graph shown above by TensorFlow.<br>\nIn this figure, the unknown variables (trainable variables) of the optimization are shown in red, pre-computable constants in blue. The inter-signal range biases are parameters shared to all times for each SignalType.</p>\n<p>Due to the availability of base station observations, only GPS, GLONASS, and Galileo were used for pseudorange observation data. For the doppler and ADR observations, I used data from all constellations, including QZSS and BeiDou.<br>\nIn addition, the weighting of residuals by RecievedSvTimeUncertaintyNanos, PseudoRangeUncertaintyMetersPerSecond, and AcuumulatedDeltaRangeUncertaintyMeters contributed to accuracy improvement.</p>\n<p>To take into account the equality constraints of the state equation, augmented Lagrangian method is used to optimize the evaluation function defined by the above computation graph.</p>\n<p><img src=\"https://user-images.githubusercontent.com/309785/181914819-7604ffec-7d19-4940-a3f0-bcd261917516.png\" alt=\"augmented_Lagrangian\"></p>\n<p>Where f(x) is the evaluation function and h(x) is the equality constraint corresponding to the above state equations.<br>\nMore specifically, the following steps are repeated alternately: <br>\n(1) minimizing the augmented Lagrangian by Adam optimizer.<br>\n(2) updating the dual variables of the equality constraints (λ).</p>\n<p>Convergence of this optimization is slow! But, it can robustly compute solutions with good accuracy.</p>\n<p>Note that double precision floating point is used in all computations to eliminate floating point rounding errors. For this reason, optimization is performed using the CPU instead of the GPU.</p>\n<h1>Workaround for pseudorange-rate of XiaomiMi8</h1>\n<p><img src=\"https://user-images.githubusercontent.com/309785/181914836-97fa4def-3d2f-4d55-9534-33fd43fc8463.png\" alt=\"velocity_XiaomiMi8\"></p>\n<p>This figure shows comparison between ground-truth velocity and estimated velocity by weighted least-squares method using pseudorange-rate in 2021-12-07-US-LAX-1/XiaomiMi8.<br>\nThis shows that there is a time discrepancy of about 0.5 seconds between the estimated velocity and the ground-truth (WHY?). Therefore, correcting for this time discrepancy improves the accuracy of position estimation.<br>\nIn my solution, this problem is addressed by a correction term (acceleration × time delta) for smartphone velocity and adding the time delta parameter to the optimization variables. (This workaround is applied to data other than XiaomiMi8.) This was quite effective in improving public LB, because there are many XiaomiMi8 in public test data.</p>\n<p><img src=\"https://user-images.githubusercontent.com/309785/181914823-1272523e-2a04-4d4a-b6d8-e0bf64797501.png\" alt=\"modify_velocity\"></p>\n<h1>Ensemble of multiple base station solutions</h1>\n<p><img src=\"https://user-images.githubusercontent.com/309785/181917131-42d34aa8-8e9a-4fa3-b9ca-c06204c6fe58.jpg\" alt=\"google-earth\"></p>\n<p>I had originally planed to interpolate solutions from multiple base stations based on the location of base stations and smartphone, but I found that the correlation between baseline length and position error was not so clear, and simply averaging the solutions from multiple base stations was slightly more accurate. The reason for this is assumed to be that the pseudorange data of the base stations also contains noise of about 0.5 to 1.0 m. Therefore, the median of the solutions of multiple base stations was used as the final submission, after correcting the difference between the ionospheric and tropospheric delays at the base stations and smartphone locations using external data and the model.</p>",
      "rawMarkdown": "Thank you to hosts for hosting this great competition. \nAt the start of this competition, I was concerned that this competition would not attract many Kagglers, because machine learning was not very effective for GNSS data in the last year competition, but I am surprised that the competition was at very high level in the end. \nI would like to publish my solution here:\n\n# Summary\n\nI followed the @taroz1461's [1st place solution from last year competition](https://www.kaggle.com/competitions/google-smartphone-decimeter-challenge/discussion/262406) and combined it with [my ideas of the 5th place solution](https://www.kaggle.com/competitions/google-smartphone-decimeter-challenge/discussion/261959).\n\nThe key point of the solution is to formulate and solve the integration of GNSS observations (pseudorange, pseudorange-rate(doppler), ADR) and trajectory smoothing as a single optimization problem. This provides following advantages: accuracy of satellite positioning due to satellite constellation is automatically taken into account in trajectory smoothing, and satellite positioning can be continued even if the number of satellites is temporarily less than required. In addition, by adding switchable constraints to all GNSS observations, satellite positioning becomes robust.\n\nOptimization was done by writing an evaluation function and a custom solver for handling equality constraints using TensorFlow (not NN). The flexibility of the deep learning framework was helpful for quickly testing ideas, but most ideas did not contribute to accuracy improvement in this competition. In the end, optimization part was fairly simple.\n\n![summary](https://user-images.githubusercontent.com/309785/181914831-1e6e3068-4046-4468-9070-dabc8d5975ef.png)\n\n# Input Data & Preprocessing\n\n+ Raw GNSS log (supplemental/gnss_log.txt): \ndevice_gnss.csv had null lines for SignalType, satellite position, etc. So I used gnss_log.txt and added columns to the obtained data-frame.\n\n+ Observation of pseudorange of base stations:\nI used [UNAVCO's](https://www.unavco.org/data/gps-gnss/gps-gnss.html) 15-second rate observation data and calculated the residual between the base station pseudorange and distance between satellite and base station. Although it is not necessary to correct for ionospheric and tropospheric delays in relative positioning, I corrected for ionospheric and tropospheric delays for both base stations and smartphone, since I also used base station data with long baseline length when performing multiple base station ensembles.\n\n+ Precise ephemeris:\nI used satellite orbit data from [GNSS IGS MGEX Products](https://cddis.nasa.gov/Data_and_Derived_Products/GNSS/gnss_mgex_products.html) to obtain the satellite's position and velocity. Although the accuracy of the satellite position is not so critical in relative positioning, it was easier to interpolate the position data from the precise ephemeris than to calculate the satellite orbit from the navigation message.\n\n+ Base station location data:\nI downloaded time series data of base station location data from [Nevada Geodetic Laboratory](http://geodesy.unr.edu/index.php), and used the average over the test data period (2021/4/28 - 2022/4/25) as the base station location.\n\n+ [Ionospheric delay data](https://cddis.nasa.gov/Data_and_Derived_Products/GNSS/atmospheric_products.html#iono):\nIt was used to compensate for the difference of ionospheric delay between the base station location and the smartphone location. The contribution to the final accuracy of using external data on ionospheric delay should be small, since the ionospheric delay should be mostly cancelled by relative positioning.\n\n+ WlsPosition in device_gnss.csv:\nI used the smoothed results of the baseline locations provided by the host to pre-compute the Sagnac effect, ionospheric delay, and tropospheric delay.\n\n+ I did not use IMU sensor data at all in this competition because sensor data is noisy and irregular. I did a lot of work on signal processing of sensor data in the last year competition, but it did not contribute much to the improvement of my score.\n\n+ Since I did not use machine learning, I used ground-truth data only to check accuracy of prediction results.\n\n# Optimization\n\n## State-space representation of smartphone trajectory\n\nThe position, velocity, acceleration, and jerk (derivative of acceleration) of the smartphone are represented in the ECEF coordinate system, and discretized state equations are created assuming that the jerk is piecewise constant.\n\n![state_equation](https://user-images.githubusercontent.com/309785/181914830-93f67930-4d52-4b62-a6b1-de88f09867cd.png)\n\nUnder this assumption, the position of the smartphone is represented by a piecewise cubic polynomial, so the position of the smartphone at any time can be accurately calculated by third-order Hermite interpolation using the position and velocity of collocation points (same for velocity).\n\nUsing the position and velocity of the smartphone calculated by this method and the position and velocity of the satellites, the distance between the satellites and the smartphone and its derivative are calculated, and the trajectory of the smartphone is calculated by minimizing the residual between them and the GNSS observation data.\n\n\n## Position smoothing via quadratic programing\n\nSmoothing is performed on the positions obtained by the weighted least squares method of psedurorange by solving a quadratic programming problem with the sum of the squares of the position residuals and jerks as the evaluation function.\nThe coefficients of the position observation matrix are calculated based on Hermite interpolation of the smartphone position based on the satellite positioning time.\n\nIn last year's competition, I added velocity and acceleration terms to the evaluation function of a similar quadratic programming problem to improve accuracy, but in this competition, I did not use doppler and ADR observations at this point because the purpose of quadratic programming problems is to obtain initial values for the global optimization that follows.\n\n## Global optimization\n\n![computation](https://user-images.githubusercontent.com/309785/181914822-b29b3c62-4973-4320-b86c-c1ddf800c332.png)\n\nThe evaluation function is calculated by the computation graph shown above by TensorFlow.\nIn this figure, the unknown variables (trainable variables) of the optimization are shown in red, pre-computable constants in blue. The inter-signal range biases are parameters shared to all times for each SignalType.\n\nDue to the availability of base station observations, only GPS, GLONASS, and Galileo were used for pseudorange observation data. For the doppler and ADR observations, I used data from all constellations, including QZSS and BeiDou.\nIn addition, the weighting of residuals by RecievedSvTimeUncertaintyNanos, PseudoRangeUncertaintyMetersPerSecond, and AcuumulatedDeltaRangeUncertaintyMeters contributed to accuracy improvement.\n\nTo take into account the equality constraints of the state equation, augmented Lagrangian method is used to optimize the evaluation function defined by the above computation graph.\n\n![augmented_Lagrangian](https://user-images.githubusercontent.com/309785/181914819-7604ffec-7d19-4940-a3f0-bcd261917516.png)\n\nWhere f(x) is the evaluation function and h(x) is the equality constraint corresponding to the above state equations.\nMore specifically, the following steps are repeated alternately: \n(1) minimizing the augmented Lagrangian by Adam optimizer.\n(2) updating the dual variables of the equality constraints (λ).\n\nConvergence of this optimization is slow! But, it can robustly compute solutions with good accuracy.\n\nNote that double precision floating point is used in all computations to eliminate floating point rounding errors. For this reason, optimization is performed using the CPU instead of the GPU.\n\n# Workaround for pseudorange-rate of XiaomiMi8\n\n![velocity_XiaomiMi8](https://user-images.githubusercontent.com/309785/181914836-97fa4def-3d2f-4d55-9534-33fd43fc8463.png)\n\nThis figure shows comparison between ground-truth velocity and estimated velocity by weighted least-squares method using pseudorange-rate in 2021-12-07-US-LAX-1/XiaomiMi8.\nThis shows that there is a time discrepancy of about 0.5 seconds between the estimated velocity and the ground-truth (WHY?). Therefore, correcting for this time discrepancy improves the accuracy of position estimation.\nIn my solution, this problem is addressed by a correction term (acceleration × time delta) for smartphone velocity and adding the time delta parameter to the optimization variables. (This workaround is applied to data other than XiaomiMi8.) This was quite effective in improving public LB, because there are many XiaomiMi8 in public test data.\n\n![modify_velocity](https://user-images.githubusercontent.com/309785/181914823-1272523e-2a04-4d4a-b6d8-e0bf64797501.png)\n\n# Ensemble of multiple base station solutions\n\n![google-earth](https://user-images.githubusercontent.com/309785/181917131-42d34aa8-8e9a-4fa3-b9ca-c06204c6fe58.jpg)\n\nI had originally planed to interpolate solutions from multiple base stations based on the location of base stations and smartphone, but I found that the correlation between baseline length and position error was not so clear, and simply averaging the solutions from multiple base stations was slightly more accurate. The reason for this is assumed to be that the pseudorange data of the base stations also contains noise of about 0.5 to 1.0 m. Therefore, the median of the solutions of multiple base stations was used as the final submission, after correcting the difference between the ionospheric and tropospheric delays at the base stations and smartphone locations using external data and the model.",
      "votes": null
    },
    {
      "id": "1877584",
      "postDate": "07/30/2022 19:55:33",
      "content": "<p>Congrats for impressive score and thanks for detailed solution description!</p>\n<p>I have couple of questions:<br>\n1) how adr improves final score comparing to only doppler velocity estimation? what was your adr loss implementation? <br>\n2) how much ensembling of base stations improved score comparing to nearest base station only?</p>",
      "rawMarkdown": "Congrats for impressive score and thanks for detailed solution description!\n\nI have couple of questions:\n1) how adr improves final score comparing to only doppler velocity estimation? what was your adr loss implementation? \n2) how much ensembling of base stations improved score comparing to nearest base station only?",
      "votes": null
    },
    {
      "id": "1877886",
      "postDate": "07/31/2022 04:11:25",
      "content": "<p>Congrats! Amazing solution.</p>",
      "rawMarkdown": "Congrats! Amazing solution.",
      "votes": null
    },
    {
      "id": "1878104",
      "postDate": "07/31/2022 07:58:14",
      "content": "<p>Congrats for 5th place, and thank you for sharing your solution. It’s beautiful approach and there is much to learn from it.<br>\nI tried global optimization approach by using PyTorch too, but didn’t work.<br>\nI have three questions about optimization by DL framework.</p>\n<ol>\n<li>How do you have the multi trainable variables? I assume that the weight and bias of the created linear layers are treated as trainable variables , as in <a href=\"https://www.kaggle.com/competitions/google-smartphone-decimeter-challenge/discussion/262355\" target=\"_blank\">last year's 3rd place</a>. However, the number of available satellites varies with utc time.  I'm interested in how you are implementing.</li>\n<li>Did you optimize for 1 iteration = 1 utc time, or 1 iteration = multiple utc times by mini-batch?</li>\n<li>How long does it take for your global optimization per one collection data?</li>\n</ol>",
      "rawMarkdown": "Congrats for 5th place, and thank you for sharing your solution. It’s beautiful approach and there is much to learn from it.\nI tried global optimization approach by using PyTorch too, but didn’t work.\nI have three questions about optimization by DL framework.\n\n1. How do you have the multi trainable variables? I assume that the weight and bias of the created linear layers are treated as trainable variables , as in [last year's 3rd place](https://www.kaggle.com/competitions/google-smartphone-decimeter-challenge/discussion/262355). However, the number of available satellites varies with utc time.  I'm interested in how you are implementing.\n2. Did you optimize for 1 iteration = 1 utc time, or 1 iteration = multiple utc times by mini-batch?\n3. How long does it take for your global optimization per one collection data?",
      "votes": null
    },
    {
      "id": "1878290",
      "postDate": "07/31/2022 10:57:58",
      "content": "<p>Congrats! </p>",
      "rawMarkdown": "Congrats!",
      "votes": null
    },
    {
      "id": "1878396",
      "postDate": "07/31/2022 11:55:02",
      "content": "<p>Thank you for your question.<br>\n1) Since I knew from the previous competition that ADR (carrier phase) is the key to high accuracy, so I did not properly examine the patterns excluding ADR. Observation data with resets and cycle slips in AccumulatedDeltaRangeState or with timestamp differences of 1.5 sec or more are ignored by multiplying the residuals by zero.<br>\nThe square of AccumulatedDeltaRangeUncertaintyMeters plus the default value (hyperparameter) is used as the variance of the ADR observation. Otherwise, the method is the same as in <a href=\"https://nikosuenderhauf.github.io/assets/papers/IROS12-switchableConstraints.pdf\" target=\"_blank\">the original paper</a>.<br>\n2) The average improvement of the ensemble was about 0.05m, but not all data showed an improvement.</p>",
      "rawMarkdown": "Thank you for your question.\n1) Since I knew from the previous competition that ADR (carrier phase) is the key to high accuracy, so I did not properly examine the patterns excluding ADR. Observation data with resets and cycle slips in AccumulatedDeltaRangeState or with timestamp differences of 1.5 sec or more are ignored by multiplying the residuals by zero.\nThe square of AccumulatedDeltaRangeUncertaintyMeters plus the default value (hyperparameter) is used as the variance of the ADR observation. Otherwise, the method is the same as in [the original paper](https://nikosuenderhauf.github.io/assets/papers/IROS12-switchableConstraints.pdf).\n2) The average improvement of the ensemble was about 0.05m, but not all data showed an improvement.",
      "votes": null
    },
    {
      "id": "1878434",
      "postDate": "07/31/2022 12:10:22",
      "content": "<p>complex but simple explanation</p>",
      "rawMarkdown": "complex but simple explanation",
      "votes": null
    },
    {
      "id": "1878530",
      "postDate": "07/31/2022 12:58:16",
      "content": "<ol>\n<li>I used Tensorflow's low-level API (e.g., tf.Variable) to create the computation graph, not Keras API. The unknown variables of position, velocity, acceleration, and jerk are equally spaced data in 1 second increments, but since the position and velocity corresponding to each satellite observation epoch are interpolated and then replicated by tf.gather, it does not matter if the number of satellites is different for each epoch.</li>\n<li>I think the answer to that question would be full-batch, since the gradient is calculated using all the data for each iteration.</li>\n<li>It takes more than 24 hours to inference 206 train and test data without ensembles, run in 2 parallel on a 16-core CPU, which would be about 15 minutes per data. Because it is better to spend more time than to truncate iterations and reduce the accuracy in csv-competition, I did not even implement a convergence criterion. The score improved when I just increased the number of iterations in the last submission, so I could have made more improvements about the optimization algorithm.</li>\n</ol>",
      "rawMarkdown": "1. I used Tensorflow's low-level API (e.g., tf.Variable) to create the computation graph, not Keras API. The unknown variables of position, velocity, acceleration, and jerk are equally spaced data in 1 second increments, but since the position and velocity corresponding to each satellite observation epoch are interpolated and then replicated by tf.gather, it does not matter if the number of satellites is different for each epoch.\n2. I think the answer to that question would be full-batch, since the gradient is calculated using all the data for each iteration.\n3. It takes more than 24 hours to inference 206 train and test data without ensembles, run in 2 parallel on a 16-core CPU, which would be about 15 minutes per data. Because it is better to spend more time than to truncate iterations and reduce the accuracy in csv-competition, I did not even implement a convergence criterion. The score improved when I just increased the number of iterations in the last submission, so I could have made more improvements about the optimization algorithm.",
      "votes": null
    },
    {
      "id": "1878660",
      "postDate": "07/31/2022 14:31:23",
      "content": "<p>Congrats and thanks for sharing all the details of your approach.</p>",
      "rawMarkdown": "Congrats and thanks for sharing all the details of your approach.",
      "votes": null
    },
    {
      "id": "1878696",
      "postDate": "07/31/2022 14:52:57",
      "content": "<p>I see, it is very helpful. I will retry my approach.<br>\nThanks for your sharing!</p>",
      "rawMarkdown": "I see, it is very helpful. I will retry my approach.\nThanks for your sharing!",
      "votes": null
    },
    {
      "id": "1880167",
      "postDate": "08/01/2022 14:02:58",
      "content": "<p>thank you sir for this</p>",
      "rawMarkdown": "thank you sir for this",
      "votes": null
    },
    {
      "id": "1880678",
      "postDate": "08/02/2022 00:15:50",
      "content": "<p>Thank you for sharing the solution!<br>\nFrom my experiments, the XiaomiMi8's Doppler agreed well with the GPS time with a time offset of 600 ms. What was the actual value of the time offset estimation result with your method? Also, was the time offset estimation effective on other smartphones?</p>",
      "rawMarkdown": "Thank you for sharing the solution!\nFrom my experiments, the XiaomiMi8's Doppler agreed well with the GPS time with a time offset of 600 ms. What was the actual value of the time offset estimation result with your method? Also, was the time offset estimation effective on other smartphones?",
      "votes": null
    },
    {
      "id": "1880703",
      "postDate": "08/02/2022 01:10:17",
      "content": "<p>Congrats on 5th place!</p>",
      "rawMarkdown": "Congrats on 5th place!",
      "votes": null
    },
    {
      "id": "1881430",
      "postDate": "08/02/2022 13:45:42",
      "content": "<p>The estimated time shift ranged from 430 to 470 ms for XiaomiMi8 and averaged 20 to 30 ms for other smartphones.<br>\nThe trend of Doppler time shift is consistent for smartphones other than XiaomiMi8, but the impact on the accuracy of position estimation will be negligible.</p>",
      "rawMarkdown": "The estimated time shift ranged from 430 to 470 ms for XiaomiMi8 and averaged 20 to 30 ms for other smartphones.\nThe trend of Doppler time shift is consistent for smartphones other than XiaomiMi8, but the impact on the accuracy of position estimation will be negligible.",
      "votes": null
    },
    {
      "id": "1881436",
      "postDate": "08/02/2022 13:52:20",
      "content": "<p>Thank you very much! I will try to verify it myself.</p>",
      "rawMarkdown": "Thank you very much! I will try to verify it myself.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1877584,
      "author_name": "kruntuid",
      "author_url": "",
      "post_date": "07/30/2022 19:55:33",
      "content": "<p>Congrats for impressive score and thanks for detailed solution description!</p>\n<p>I have couple of questions:<br>\n1) how adr improves final score comparing to only doppler velocity estimation? what was your adr loss implementation? <br>\n2) how much ensembling of base stations improved score comparing to nearest base station only?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1878396,
          "author_name": "saitodevel01",
          "author_url": "",
          "post_date": "07/31/2022 11:55:02",
          "content": "<p>Thank you for your question.<br>\n1) Since I knew from the previous competition that ADR (carrier phase) is the key to high accuracy, so I did not properly examine the patterns excluding ADR. Observation data with resets and cycle slips in AccumulatedDeltaRangeState or with timestamp differences of 1.5 sec or more are ignored by multiplying the residuals by zero.<br>\nThe square of AccumulatedDeltaRangeUncertaintyMeters plus the default value (hyperparameter) is used as the variance of the ADR observation. Otherwise, the method is the same as in <a href=\"https://nikosuenderhauf.github.io/assets/papers/IROS12-switchableConstraints.pdf\" target=\"_blank\">the original paper</a>.<br>\n2) The average improvement of the ensemble was about 0.05m, but not all data showed an improvement.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1877886,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "07/31/2022 04:11:25",
      "content": "<p>Congrats! Amazing solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1878104,
      "author_name": "kuto0633",
      "author_url": "",
      "post_date": "07/31/2022 07:58:14",
      "content": "<p>Congrats for 5th place, and thank you for sharing your solution. It’s beautiful approach and there is much to learn from it.<br>\nI tried global optimization approach by using PyTorch too, but didn’t work.<br>\nI have three questions about optimization by DL framework.</p>\n<ol>\n<li>How do you have the multi trainable variables? I assume that the weight and bias of the created linear layers are treated as trainable variables , as in <a href=\"https://www.kaggle.com/competitions/google-smartphone-decimeter-challenge/discussion/262355\" target=\"_blank\">last year's 3rd place</a>. However, the number of available satellites varies with utc time.  I'm interested in how you are implementing.</li>\n<li>Did you optimize for 1 iteration = 1 utc time, or 1 iteration = multiple utc times by mini-batch?</li>\n<li>How long does it take for your global optimization per one collection data?</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 1878530,
          "author_name": "saitodevel01",
          "author_url": "",
          "post_date": "07/31/2022 12:58:16",
          "content": "<ol>\n<li>I used Tensorflow's low-level API (e.g., tf.Variable) to create the computation graph, not Keras API. The unknown variables of position, velocity, acceleration, and jerk are equally spaced data in 1 second increments, but since the position and velocity corresponding to each satellite observation epoch are interpolated and then replicated by tf.gather, it does not matter if the number of satellites is different for each epoch.</li>\n<li>I think the answer to that question would be full-batch, since the gradient is calculated using all the data for each iteration.</li>\n<li>It takes more than 24 hours to inference 206 train and test data without ensembles, run in 2 parallel on a 16-core CPU, which would be about 15 minutes per data. Because it is better to spend more time than to truncate iterations and reduce the accuracy in csv-competition, I did not even implement a convergence criterion. The score improved when I just increased the number of iterations in the last submission, so I could have made more improvements about the optimization algorithm.</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1878696,
          "author_name": "kuto0633",
          "author_url": "",
          "post_date": "07/31/2022 14:52:57",
          "content": "<p>I see, it is very helpful. I will retry my approach.<br>\nThanks for your sharing!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1878290,
      "author_name": "flops57",
      "author_url": "",
      "post_date": "07/31/2022 10:57:58",
      "content": "<p>Congrats! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1878434,
      "author_name": "burakkaya35",
      "author_url": "",
      "post_date": "07/31/2022 12:10:22",
      "content": "<p>complex but simple explanation</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1878660,
      "author_name": "oscarm524",
      "author_url": "",
      "post_date": "07/31/2022 14:31:23",
      "content": "<p>Congrats and thanks for sharing all the details of your approach.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1880167,
      "author_name": "bilalsuppal",
      "author_url": "",
      "post_date": "08/01/2022 14:02:58",
      "content": "<p>thank you sir for this</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1880678,
      "author_name": "taroz1461",
      "author_url": "",
      "post_date": "08/02/2022 00:15:50",
      "content": "<p>Thank you for sharing the solution!<br>\nFrom my experiments, the XiaomiMi8's Doppler agreed well with the GPS time with a time offset of 600 ms. What was the actual value of the time offset estimation result with your method? Also, was the time offset estimation effective on other smartphones?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1881430,
          "author_name": "saitodevel01",
          "author_url": "",
          "post_date": "08/02/2022 13:45:42",
          "content": "<p>The estimated time shift ranged from 430 to 470 ms for XiaomiMi8 and averaged 20 to 30 ms for other smartphones.<br>\nThe trend of Doppler time shift is consistent for smartphones other than XiaomiMi8, but the impact on the accuracy of position estimation will be negligible.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1881436,
          "author_name": "taroz1461",
          "author_url": "",
          "post_date": "08/02/2022 13:52:20",
          "content": "<p>Thank you very much! I will try to verify it myself.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1880703,
      "author_name": "davidtoth77",
      "author_url": "",
      "post_date": "08/02/2022 01:10:17",
      "content": "<p>Congrats on 5th place!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1877297": "Thank you to hosts for hosting this great competition. \nAt the start of this competition, I was concerned that this competition would not attract many Kagglers, because machine learning was not very effective for GNSS data in the last year competition, but I am surprised that the competition was at very high level in the end. \nI would like to publish my solution here:\n\n# Summary\n\nI followed the @taroz1461's [1st place solution from last year competition](https://www.kaggle.com/competitions/google-smartphone-decimeter-challenge/discussion/262406) and combined it with [my ideas of the 5th place solution](https://www.kaggle.com/competitions/google-smartphone-decimeter-challenge/discussion/261959).\n\nThe key point of the solution is to formulate and solve the integration of GNSS observations (pseudorange, pseudorange-rate(doppler), ADR) and trajectory smoothing as a single optimization problem. This provides following advantages: accuracy of satellite positioning due to satellite constellation is automatically taken into account in trajectory smoothing, and satellite positioning can be continued even if the number of satellites is temporarily less than required. In addition, by adding switchable constraints to all GNSS observations, satellite positioning becomes robust.\n\nOptimization was done by writing an evaluation function and a custom solver for handling equality constraints using TensorFlow (not NN). The flexibility of the deep learning framework was helpful for quickly testing ideas, but most ideas did not contribute to accuracy improvement in this competition. In the end, optimization part was fairly simple.\n\n![summary](https://user-images.githubusercontent.com/309785/181914831-1e6e3068-4046-4468-9070-dabc8d5975ef.png)\n\n# Input Data & Preprocessing\n\n+ Raw GNSS log (supplemental/gnss_log.txt): \ndevice_gnss.csv had null lines for SignalType, satellite position, etc. So I used gnss_log.txt and added columns to the obtained data-frame.\n\n+ Observation of pseudorange of base stations:\nI used [UNAVCO's](https://www.unavco.org/data/gps-gnss/gps-gnss.html) 15-second rate observation data and calculated the residual between the base station pseudorange and distance between satellite and base station. Although it is not necessary to correct for ionospheric and tropospheric delays in relative positioning, I corrected for ionospheric and tropospheric delays for both base stations and smartphone, since I also used base station data with long baseline length when performing multiple base station ensembles.\n\n+ Precise ephemeris:\nI used satellite orbit data from [GNSS IGS MGEX Products](https://cddis.nasa.gov/Data_and_Derived_Products/GNSS/gnss_mgex_products.html) to obtain the satellite's position and velocity. Although the accuracy of the satellite position is not so critical in relative positioning, it was easier to interpolate the position data from the precise ephemeris than to calculate the satellite orbit from the navigation message.\n\n+ Base station location data:\nI downloaded time series data of base station location data from [Nevada Geodetic Laboratory](http://geodesy.unr.edu/index.php), and used the average over the test data period (2021/4/28 - 2022/4/25) as the base station location.\n\n+ [Ionospheric delay data](https://cddis.nasa.gov/Data_and_Derived_Products/GNSS/atmospheric_products.html#iono):\nIt was used to compensate for the difference of ionospheric delay between the base station location and the smartphone location. The contribution to the final accuracy of using external data on ionospheric delay should be small, since the ionospheric delay should be mostly cancelled by relative positioning.\n\n+ WlsPosition in device_gnss.csv:\nI used the smoothed results of the baseline locations provided by the host to pre-compute the Sagnac effect, ionospheric delay, and tropospheric delay.\n\n+ I did not use IMU sensor data at all in this competition because sensor data is noisy and irregular. I did a lot of work on signal processing of sensor data in the last year competition, but it did not contribute much to the improvement of my score.\n\n+ Since I did not use machine learning, I used ground-truth data only to check accuracy of prediction results.\n\n# Optimization\n\n## State-space representation of smartphone trajectory\n\nThe position, velocity, acceleration, and jerk (derivative of acceleration) of the smartphone are represented in the ECEF coordinate system, and discretized state equations are created assuming that the jerk is piecewise constant.\n\n![state_equation](https://user-images.githubusercontent.com/309785/181914830-93f67930-4d52-4b62-a6b1-de88f09867cd.png)\n\nUnder this assumption, the position of the smartphone is represented by a piecewise cubic polynomial, so the position of the smartphone at any time can be accurately calculated by third-order Hermite interpolation using the position and velocity of collocation points (same for velocity).\n\nUsing the position and velocity of the smartphone calculated by this method and the position and velocity of the satellites, the distance between the satellites and the smartphone and its derivative are calculated, and the trajectory of the smartphone is calculated by minimizing the residual between them and the GNSS observation data.\n\n\n## Position smoothing via quadratic programing\n\nSmoothing is performed on the positions obtained by the weighted least squares method of psedurorange by solving a quadratic programming problem with the sum of the squares of the position residuals and jerks as the evaluation function.\nThe coefficients of the position observation matrix are calculated based on Hermite interpolation of the smartphone position based on the satellite positioning time.\n\nIn last year's competition, I added velocity and acceleration terms to the evaluation function of a similar quadratic programming problem to improve accuracy, but in this competition, I did not use doppler and ADR observations at this point because the purpose of quadratic programming problems is to obtain initial values for the global optimization that follows.\n\n## Global optimization\n\n![computation](https://user-images.githubusercontent.com/309785/181914822-b29b3c62-4973-4320-b86c-c1ddf800c332.png)\n\nThe evaluation function is calculated by the computation graph shown above by TensorFlow.\nIn this figure, the unknown variables (trainable variables) of the optimization are shown in red, pre-computable constants in blue. The inter-signal range biases are parameters shared to all times for each SignalType.\n\nDue to the availability of base station observations, only GPS, GLONASS, and Galileo were used for pseudorange observation data. For the doppler and ADR observations, I used data from all constellations, including QZSS and BeiDou.\nIn addition, the weighting of residuals by RecievedSvTimeUncertaintyNanos, PseudoRangeUncertaintyMetersPerSecond, and AcuumulatedDeltaRangeUncertaintyMeters contributed to accuracy improvement.\n\nTo take into account the equality constraints of the state equation, augmented Lagrangian method is used to optimize the evaluation function defined by the above computation graph.\n\n![augmented_Lagrangian](https://user-images.githubusercontent.com/309785/181914819-7604ffec-7d19-4940-a3f0-bcd261917516.png)\n\nWhere f(x) is the evaluation function and h(x) is the equality constraint corresponding to the above state equations.\nMore specifically, the following steps are repeated alternately: \n(1) minimizing the augmented Lagrangian by Adam optimizer.\n(2) updating the dual variables of the equality constraints (λ).\n\nConvergence of this optimization is slow! But, it can robustly compute solutions with good accuracy.\n\nNote that double precision floating point is used in all computations to eliminate floating point rounding errors. For this reason, optimization is performed using the CPU instead of the GPU.\n\n# Workaround for pseudorange-rate of XiaomiMi8\n\n![velocity_XiaomiMi8](https://user-images.githubusercontent.com/309785/181914836-97fa4def-3d2f-4d55-9534-33fd43fc8463.png)\n\nThis figure shows comparison between ground-truth velocity and estimated velocity by weighted least-squares method using pseudorange-rate in 2021-12-07-US-LAX-1/XiaomiMi8.\nThis shows that there is a time discrepancy of about 0.5 seconds between the estimated velocity and the ground-truth (WHY?). Therefore, correcting for this time discrepancy improves the accuracy of position estimation.\nIn my solution, this problem is addressed by a correction term (acceleration × time delta) for smartphone velocity and adding the time delta parameter to the optimization variables. (This workaround is applied to data other than XiaomiMi8.) This was quite effective in improving public LB, because there are many XiaomiMi8 in public test data.\n\n![modify_velocity](https://user-images.githubusercontent.com/309785/181914823-1272523e-2a04-4d4a-b6d8-e0bf64797501.png)\n\n# Ensemble of multiple base station solutions\n\n![google-earth](https://user-images.githubusercontent.com/309785/181917131-42d34aa8-8e9a-4fa3-b9ca-c06204c6fe58.jpg)\n\nI had originally planed to interpolate solutions from multiple base stations based on the location of base stations and smartphone, but I found that the correlation between baseline length and position error was not so clear, and simply averaging the solutions from multiple base stations was slightly more accurate. The reason for this is assumed to be that the pseudorange data of the base stations also contains noise of about 0.5 to 1.0 m. Therefore, the median of the solutions of multiple base stations was used as the final submission, after correcting the difference between the ionospheric and tropospheric delays at the base stations and smartphone locations using external data and the model.",
    "1877584": "Congrats for impressive score and thanks for detailed solution description!\n\nI have couple of questions:\n1) how adr improves final score comparing to only doppler velocity estimation? what was your adr loss implementation? \n2) how much ensembling of base stations improved score comparing to nearest base station only?",
    "1877886": "Congrats! Amazing solution.",
    "1878104": "Congrats for 5th place, and thank you for sharing your solution. It’s beautiful approach and there is much to learn from it.\nI tried global optimization approach by using PyTorch too, but didn’t work.\nI have three questions about optimization by DL framework.\n\n1. How do you have the multi trainable variables? I assume that the weight and bias of the created linear layers are treated as trainable variables , as in [last year's 3rd place](https://www.kaggle.com/competitions/google-smartphone-decimeter-challenge/discussion/262355). However, the number of available satellites varies with utc time.  I'm interested in how you are implementing.\n2. Did you optimize for 1 iteration = 1 utc time, or 1 iteration = multiple utc times by mini-batch?\n3. How long does it take for your global optimization per one collection data?",
    "1878290": "Congrats!",
    "1878396": "Thank you for your question.\n1) Since I knew from the previous competition that ADR (carrier phase) is the key to high accuracy, so I did not properly examine the patterns excluding ADR. Observation data with resets and cycle slips in AccumulatedDeltaRangeState or with timestamp differences of 1.5 sec or more are ignored by multiplying the residuals by zero.\nThe square of AccumulatedDeltaRangeUncertaintyMeters plus the default value (hyperparameter) is used as the variance of the ADR observation. Otherwise, the method is the same as in [the original paper](https://nikosuenderhauf.github.io/assets/papers/IROS12-switchableConstraints.pdf).\n2) The average improvement of the ensemble was about 0.05m, but not all data showed an improvement.",
    "1878434": "complex but simple explanation",
    "1878530": "1. I used Tensorflow's low-level API (e.g., tf.Variable) to create the computation graph, not Keras API. The unknown variables of position, velocity, acceleration, and jerk are equally spaced data in 1 second increments, but since the position and velocity corresponding to each satellite observation epoch are interpolated and then replicated by tf.gather, it does not matter if the number of satellites is different for each epoch.\n2. I think the answer to that question would be full-batch, since the gradient is calculated using all the data for each iteration.\n3. It takes more than 24 hours to inference 206 train and test data without ensembles, run in 2 parallel on a 16-core CPU, which would be about 15 minutes per data. Because it is better to spend more time than to truncate iterations and reduce the accuracy in csv-competition, I did not even implement a convergence criterion. The score improved when I just increased the number of iterations in the last submission, so I could have made more improvements about the optimization algorithm.",
    "1878660": "Congrats and thanks for sharing all the details of your approach.",
    "1878696": "I see, it is very helpful. I will retry my approach.\nThanks for your sharing!",
    "1880167": "thank you sir for this",
    "1880678": "Thank you for sharing the solution!\nFrom my experiments, the XiaomiMi8's Doppler agreed well with the GPS time with a time offset of 600 ms. What was the actual value of the time offset estimation result with your method? Also, was the time offset estimation effective on other smartphones?",
    "1880703": "Congrats on 5th place!",
    "1881430": "The estimated time shift ranged from 430 to 470 ms for XiaomiMi8 and averaged 20 to 30 ms for other smartphones.\nThe trend of Doppler time shift is consistent for smartphones other than XiaomiMi8, but the impact on the accuracy of position estimation will be negligible.",
    "1881436": "Thank you very much! I will try to verify it myself."
  },
  "source": "meta"
}