{
  "id": 340663,
  "title": "[43rd place] RTKLIB & Lots of Postprocessing",
  "url": "/competitions/smartphone-decimeter-2022/writeups/john-mitchell-43rd-place-rtklib-lots-of-postproces",
  "author_name": "",
  "post_date": "2022-08-03T14:01:33.657Z",
  "votes": 10,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Thanks to the hosts and to eveyone who took part and contributed to this very interesting and challenging competition. I am very much looking forward to reading through write-ups from the gold medal teams. As you will gather from what follows, I concentrated on firstly getting RTKLIB to work for me, and secondly on adding layers of postprocessing.</p>\n<p><strong>Step 0.  My starting point was a slightly adjusted version of <a href=\"https://www.kaggle.com/taroz1461\" target=\"_blank\">@taroz1461</a>’s notebook</strong></p>\n<p><a href=\"https://www.kaggle.com/code/taroz1461/carrier-smoothing-robust-wls-kalman-smoother\" target=\"_blank\">Carrier Smoothing + Robust WLS + Kalman Smoother</a></p>\n<p>Public score 2.964, Private score 3.462</p>\n<p>Note that I clearly had overfitted this to the public LB, as my version was worse than the original on the private set.</p>\n<p>Subsequent steps are listed in logical rather than temporal order.</p>\n<p><strong>Step 1. RTKLIB.</strong></p>\n<p><a href=\"https://www.kaggle.com/timeverett\" target=\"_blank\">@timeverett</a> provided a notebook</p>\n<p><a href=\"https://www.kaggle.com/code/timeverett/getting-started-with-rtklib\" target=\"_blank\">Getting Started with RTKLIB</a></p>\n<p>giving detailed instructions on the installation of RKTLIB. From the quoted LB scores, this offered an improvement of close to 0.3 m on the public LB and more than 1.0 m on the private set,  more than the cumulative effect of the other steps of my workflow. Nonetheless, it was far from straightforward to get the software working correctly in my Windows environment. Most of my issues arose from the code not liking Windows paths. My initial problem turned out to be that I'd tried to run the code from a Windows folder with a space in its name, the space arising from the formatting of my user name, which threw an error. I therefore created a second instance, this time in a folder with no space in the name. The python code needed to be told explicitly about Windows paths. Thus, in step 4 of the process defined in <a href=\"https://www.kaggle.com/timeverett\" target=\"_blank\">@timeverett</a>'s notebook (so step 4 of step 1 in my nomenclature - sorry!), I changed the python code to explicitly add the required paths in the correct OS format. After three separate days of struggle, I finally got the code to run correctly.</p>\n<p>Public score 2.682, Private score 2.189</p>\n<p><strong>Step 2. Generic Bias Correction</strong></p>\n<p>The excellent notebooks</p>\n<p><a href=\"https://www.kaggle.com/code/saitodevel01/gsdc-bias-eda\" target=\"_blank\">GSDC - Bias EDA</a></p>\n<p>and</p>\n<p><a href=\"https://www.kaggle.com/code/saitodevel01/gsdc-bias-correction\" target=\"_blank\">GSDC - Bias Correction</a></p>\n<p>by <a href=\"https://www.kaggle.com/saitodevel01\" target=\"_blank\">@saitodevel01</a>, written for the previous GSDC competition in 2021, demonstrate how the positions of the phone and driver in the car create biases between the recorded and actual positions in both the front-back and left-right directions.<br>\nIn outline, the phone has better sightlines to satellites located in front of the car, since it is placed close to the to front windscreen, whereas sightlines of signals from satellites located behind the car may be compromised. Similarly, the driver sits to the left of the phone (these data came from the USA) and may block signals from satellites located to the left, whereas those to the right should have better communication. This would be expected to lead to an artefactual error locating the phone too far forward and too far to the right; hence the anticipated correction would move the phone backwards and leftwards.</p>\n<p>I analysed <a href=\"https://www.kaggle.com/saitodevel01\" target=\"_blank\">@saitodevel01</a>'s work from the 2021 competition, here is an example plot showing <a href=\"https://www.kaggle.com/saitodevel01\" target=\"_blank\">@saitodevel01</a>'s latitude correction against the cosine of the angle between the car's current direction and due north:</p>\n<p><img src=\"https://i.imgur.com/HjF7Nmh.png\"></p>\n<p>The corresponding plot for longitude looks broadly similar in form. My simplification is to treat the ellipse as the signal I want to model, and ignore all other points. That leads to a constant \"lever arm\" type correction of these biases, which I express as simple trigonometric functions of the direction the car is pointing in any one time step. Thus, I can compute latitude and longitude corrections for each predicted position.</p>\n<p>My analysis of the prior work suggests that the predicted uniform generic correction is:</p>\n<p>40 cm from front to back; 27 cm from right to left.</p>\n<p>I added these corrections to the originally predicted latitude and longitude to generate a trial postprocessed submission.<br>\nHowever, some trial and error showed that the best public LB scores came from a somewhat larger front-back correction of about 1.3 times that previously suggested. This works out as about 52 cm front to back correction.</p>\n<p>For the left-right correction, surprisingly the best empirical public LB results came from a small correction of about 8 cm from left to right, that is in the counterintuitive direction. However, for simplicity, I set this generic left-right correction to zero.</p>\n<p>Public score 2.595, Private score 2.158</p>\n<p><strong>Step 3: Phone-Specific Bias Correction</strong></p>\n<p>Some consideration of possible relevant effects led me to consider whether the bias correction might depend on the model of phone. There are various possible explanations for this. Possibly some aspect of the phone's hardware or software, or alternatively some quirk of the individual devices used in the experiment, meant that some phones were better able than others to communicate with satellites via compromised sightlines or through atmospheric or environmental interference. Hence, corrections could vary between phones. Alternatively, if each phone has a designated holder or position in the car, its results may be affected by sightlines and/or lever arm correction errors related to that position. For example. one phone holder may have better views backwards or to the left than another, depending on the extent to which the driver, seats, or experimental equipment block the view.  This seemed worth investigating, so a brief analysis was made using training set data, giving these results:</p>\n<p>GooglePixel4    <br>\nForward error   -0.1798<br>\nRelative to dataset -0.1639<br>\nLeft error  -0.3383<br>\nRelative to dataset -0.0136</p>\n<p>GooglePixel4XL  <br>\nForward error   0.1375<br>\nRelative to dataset 0.1535<br>\nLeft error  -0.2301<br>\nRelative to dataset 0.0947</p>\n<p>GooglePixel5    <br>\nForward error   -0.7068<br>\nRelative to dataset -0.6909<br>\nLeft error  0.2203<br>\nRelative to dataset 0.5450</p>\n<p>SamsungGalaxyS20Ultra   <br>\nForward error   0.1469<br>\nRelative to dataset 0.1628<br>\nLeft error  -1.0847<br>\nRelative to dataset -0.7599</p>\n<p>XiaomiMi8   <br>\nForward error   0.4559<br>\nRelative to dataset 0.4718<br>\nLeft error  -1.0740<br>\nRelative to dataset -0.7493</p>\n<p>It seemed that it was worth trying a phone-specific set of corrections. Firstly, I looked at the front-back corrections. I only accept these corrections if they are in the same direction as the training results, and (for the front-back correction) no larger than the original generic correction. We can do this for the four phone models that are common to the training and test data; GooglePixel4, GooglePixel5, XiaomiMi8 and SamsungGalaxyS20Ultra. There is some trial and error involved here; the two (direction and size) constraints mentioned above are designed to limit public-score chasing. I also implemented phone-specific left-right corrections.</p>\n<p>This is implemented in the notebook</p>\n<p><a href=\"https://www.kaggle.com/code/jbomitchell/phone-specific-bias-corrections-clipped\" target=\"_blank\">Phone-specific Bias Corrections (Clipped)</a></p>\n<p>Public score 2.508, Private score 2.094</p>\n<p><strong>Step 4: Stop averaging.</strong></p>\n<p>Any car journey is likely to involve periods when the vehicle is not going anywhere, whether parked up at the beginning or end, stopped at a light, or waiting for a junction or road to become clear. When we look at the data, we see substantial sequences of points where the position only changes slightly, these drifts of a few cm generally being artefacts. Despite the car being stationary, the GNSS position still appears to wander around a little, which leads to the idea of averaging coordinates over sets of points where the car was stopped. A threshold of 1 m/s seemed to work fairly well for many of my predictions, with most submission scores being improved. Some, but not all, submissions benefitted from pasting the output of the stop averaging process directly back into another iteration of the same computation. However, the RTKLIB solutions could not be improved by stop averaging. </p>\n<p>Public and private scores were not improved for my best submissions.</p>\n<p><strong>Step 5: Tectonic correction.</strong></p>\n<p>I implemented corrections for the change in location of the base stations due to tectonic movement. The relevant base stations were: SLAC for Bay Area trips, with ITRF2014 velocity:<br>\nnorthward = 0.0164 m/yr, eastward = -0.0343 m/yr, upward = 0.0009 m/yr<br>\nStation VDCY for Los Angeles trips, with ITRF 2014 velocity:<br>\nnorthward = = 0.0131 m/yr, eastward = -0.0374 m/yr, upward = 0.0009 m/yr<br>\nThe RTKLIB calculations underlying my submission used a single set of base station positions computed to be correct midway through the time period over which the data were being collected, VDCY being selected for Los Angeles trips and SLAC for Bay Area trips. The distance through which I need to shift the predication is proportional to the (signed) number of days before or after the reference date that the run takes place (N), with the annual position shift being multiplied by (N/365.25) and assigned the appropriate sign depending on the direction of the time shift.<br>\nI didn’t end up choosing a submission including this correction, based on the public score, but in fact it would have given a small improvement on the private set. The notebook is:</p>\n<p><a href=\"https://www.kaggle.com/code/jbomitchell/tectonic-correction\" target=\"_blank\">Tectonic Correction</a></p>\n<p>Public score 2.508, Private score 2.093</p>\n<p><strong>Step 6. Outlier correction.</strong></p>\n<p>In principle, it sounds straightforward to make a simple projection or projections of where a point should be predicted to lie, and compare with the actual prediction. There’s a good notebook<br>\n<a href=\"https://www.kaggle.com/code/ravishah1/gsdc2-savgol-filter-outlier-removal-with-bo\" target=\"_blank\">GSDC2 - Savgol Filter + Outlier Removal with BO</a><br>\nthat implements an algorithm of this kind. However, I found no improvements from applying this kind of approach to what I’d consider “reasonably good” submissions. Even when I focused on fixing just the most egregious 100 or so outliers, I still made my score worse. I suspect it is hard to do what the competition metric requires, which is essentially to fix a “worst 5%” prediction well enough to move it out of the highest error 95%. I had some further ideas as to how to get this to work, but with the likely payoff rather small, I did not pursue them.</p>\n<p>Public score was not improved. One of my outlier correction models would marginally have improved the private score.</p>",
  "messages": [
    {
      "id": "1877095",
      "postDate": "07/30/2022 10:50:41",
      "content": "<p>Thanks to the hosts and to eveyone who took part and contributed to this very interesting and challenging competition. I am very much looking forward to reading through write-ups from the gold medal teams. As you will gather from what follows, I concentrated on firstly getting RTKLIB to work for me, and secondly on adding layers of postprocessing.</p>\n<p><strong>Step 0.  My starting point was a slightly adjusted version of <a href=\"https://www.kaggle.com/taroz1461\" target=\"_blank\">@taroz1461</a>’s notebook</strong></p>\n<p><a href=\"https://www.kaggle.com/code/taroz1461/carrier-smoothing-robust-wls-kalman-smoother\" target=\"_blank\">Carrier Smoothing + Robust WLS + Kalman Smoother</a></p>\n<p>Public score 2.964, Private score 3.462</p>\n<p>Note that I clearly had overfitted this to the public LB, as my version was worse than the original on the private set.</p>\n<p>Subsequent steps are listed in logical rather than temporal order.</p>\n<p><strong>Step 1. RTKLIB.</strong></p>\n<p><a href=\"https://www.kaggle.com/timeverett\" target=\"_blank\">@timeverett</a> provided a notebook</p>\n<p><a href=\"https://www.kaggle.com/code/timeverett/getting-started-with-rtklib\" target=\"_blank\">Getting Started with RTKLIB</a></p>\n<p>giving detailed instructions on the installation of RKTLIB. From the quoted LB scores, this offered an improvement of close to 0.3 m on the public LB and more than 1.0 m on the private set,  more than the cumulative effect of the other steps of my workflow. Nonetheless, it was far from straightforward to get the software working correctly in my Windows environment. Most of my issues arose from the code not liking Windows paths. My initial problem turned out to be that I'd tried to run the code from a Windows folder with a space in its name, the space arising from the formatting of my user name, which threw an error. I therefore created a second instance, this time in a folder with no space in the name. The python code needed to be told explicitly about Windows paths. Thus, in step 4 of the process defined in <a href=\"https://www.kaggle.com/timeverett\" target=\"_blank\">@timeverett</a>'s notebook (so step 4 of step 1 in my nomenclature - sorry!), I changed the python code to explicitly add the required paths in the correct OS format. After three separate days of struggle, I finally got the code to run correctly.</p>\n<p>Public score 2.682, Private score 2.189</p>\n<p><strong>Step 2. Generic Bias Correction</strong></p>\n<p>The excellent notebooks</p>\n<p><a href=\"https://www.kaggle.com/code/saitodevel01/gsdc-bias-eda\" target=\"_blank\">GSDC - Bias EDA</a></p>\n<p>and</p>\n<p><a href=\"https://www.kaggle.com/code/saitodevel01/gsdc-bias-correction\" target=\"_blank\">GSDC - Bias Correction</a></p>\n<p>by <a href=\"https://www.kaggle.com/saitodevel01\" target=\"_blank\">@saitodevel01</a>, written for the previous GSDC competition in 2021, demonstrate how the positions of the phone and driver in the car create biases between the recorded and actual positions in both the front-back and left-right directions.<br>\nIn outline, the phone has better sightlines to satellites located in front of the car, since it is placed close to the to front windscreen, whereas sightlines of signals from satellites located behind the car may be compromised. Similarly, the driver sits to the left of the phone (these data came from the USA) and may block signals from satellites located to the left, whereas those to the right should have better communication. This would be expected to lead to an artefactual error locating the phone too far forward and too far to the right; hence the anticipated correction would move the phone backwards and leftwards.</p>\n<p>I analysed <a href=\"https://www.kaggle.com/saitodevel01\" target=\"_blank\">@saitodevel01</a>'s work from the 2021 competition, here is an example plot showing <a href=\"https://www.kaggle.com/saitodevel01\" target=\"_blank\">@saitodevel01</a>'s latitude correction against the cosine of the angle between the car's current direction and due north:</p>\n<p><img src=\"https://i.imgur.com/HjF7Nmh.png\"></p>\n<p>The corresponding plot for longitude looks broadly similar in form. My simplification is to treat the ellipse as the signal I want to model, and ignore all other points. That leads to a constant \"lever arm\" type correction of these biases, which I express as simple trigonometric functions of the direction the car is pointing in any one time step. Thus, I can compute latitude and longitude corrections for each predicted position.</p>\n<p>My analysis of the prior work suggests that the predicted uniform generic correction is:</p>\n<p>40 cm from front to back; 27 cm from right to left.</p>\n<p>I added these corrections to the originally predicted latitude and longitude to generate a trial postprocessed submission.<br>\nHowever, some trial and error showed that the best public LB scores came from a somewhat larger front-back correction of about 1.3 times that previously suggested. This works out as about 52 cm front to back correction.</p>\n<p>For the left-right correction, surprisingly the best empirical public LB results came from a small correction of about 8 cm from left to right, that is in the counterintuitive direction. However, for simplicity, I set this generic left-right correction to zero.</p>\n<p>Public score 2.595, Private score 2.158</p>\n<p><strong>Step 3: Phone-Specific Bias Correction</strong></p>\n<p>Some consideration of possible relevant effects led me to consider whether the bias correction might depend on the model of phone. There are various possible explanations for this. Possibly some aspect of the phone's hardware or software, or alternatively some quirk of the individual devices used in the experiment, meant that some phones were better able than others to communicate with satellites via compromised sightlines or through atmospheric or environmental interference. Hence, corrections could vary between phones. Alternatively, if each phone has a designated holder or position in the car, its results may be affected by sightlines and/or lever arm correction errors related to that position. For example. one phone holder may have better views backwards or to the left than another, depending on the extent to which the driver, seats, or experimental equipment block the view.  This seemed worth investigating, so a brief analysis was made using training set data, giving these results:</p>\n<p>GooglePixel4    <br>\nForward error   -0.1798<br>\nRelative to dataset -0.1639<br>\nLeft error  -0.3383<br>\nRelative to dataset -0.0136</p>\n<p>GooglePixel4XL  <br>\nForward error   0.1375<br>\nRelative to dataset 0.1535<br>\nLeft error  -0.2301<br>\nRelative to dataset 0.0947</p>\n<p>GooglePixel5    <br>\nForward error   -0.7068<br>\nRelative to dataset -0.6909<br>\nLeft error  0.2203<br>\nRelative to dataset 0.5450</p>\n<p>SamsungGalaxyS20Ultra   <br>\nForward error   0.1469<br>\nRelative to dataset 0.1628<br>\nLeft error  -1.0847<br>\nRelative to dataset -0.7599</p>\n<p>XiaomiMi8   <br>\nForward error   0.4559<br>\nRelative to dataset 0.4718<br>\nLeft error  -1.0740<br>\nRelative to dataset -0.7493</p>\n<p>It seemed that it was worth trying a phone-specific set of corrections. Firstly, I looked at the front-back corrections. I only accept these corrections if they are in the same direction as the training results, and (for the front-back correction) no larger than the original generic correction. We can do this for the four phone models that are common to the training and test data; GooglePixel4, GooglePixel5, XiaomiMi8 and SamsungGalaxyS20Ultra. There is some trial and error involved here; the two (direction and size) constraints mentioned above are designed to limit public-score chasing. I also implemented phone-specific left-right corrections.</p>\n<p>This is implemented in the notebook</p>\n<p><a href=\"https://www.kaggle.com/code/jbomitchell/phone-specific-bias-corrections-clipped\" target=\"_blank\">Phone-specific Bias Corrections (Clipped)</a></p>\n<p>Public score 2.508, Private score 2.094</p>\n<p><strong>Step 4: Stop averaging.</strong></p>\n<p>Any car journey is likely to involve periods when the vehicle is not going anywhere, whether parked up at the beginning or end, stopped at a light, or waiting for a junction or road to become clear. When we look at the data, we see substantial sequences of points where the position only changes slightly, these drifts of a few cm generally being artefacts. Despite the car being stationary, the GNSS position still appears to wander around a little, which leads to the idea of averaging coordinates over sets of points where the car was stopped. A threshold of 1 m/s seemed to work fairly well for many of my predictions, with most submission scores being improved. Some, but not all, submissions benefitted from pasting the output of the stop averaging process directly back into another iteration of the same computation. However, the RTKLIB solutions could not be improved by stop averaging. </p>\n<p>Public and private scores were not improved for my best submissions.</p>\n<p><strong>Step 5: Tectonic correction.</strong></p>\n<p>I implemented corrections for the change in location of the base stations due to tectonic movement. The relevant base stations were: SLAC for Bay Area trips, with ITRF2014 velocity:<br>\nnorthward = 0.0164 m/yr, eastward = -0.0343 m/yr, upward = 0.0009 m/yr<br>\nStation VDCY for Los Angeles trips, with ITRF 2014 velocity:<br>\nnorthward = = 0.0131 m/yr, eastward = -0.0374 m/yr, upward = 0.0009 m/yr<br>\nThe RTKLIB calculations underlying my submission used a single set of base station positions computed to be correct midway through the time period over which the data were being collected, VDCY being selected for Los Angeles trips and SLAC for Bay Area trips. The distance through which I need to shift the predication is proportional to the (signed) number of days before or after the reference date that the run takes place (N), with the annual position shift being multiplied by (N/365.25) and assigned the appropriate sign depending on the direction of the time shift.<br>\nI didn’t end up choosing a submission including this correction, based on the public score, but in fact it would have given a small improvement on the private set. The notebook is:</p>\n<p><a href=\"https://www.kaggle.com/code/jbomitchell/tectonic-correction\" target=\"_blank\">Tectonic Correction</a></p>\n<p>Public score 2.508, Private score 2.093</p>\n<p><strong>Step 6. Outlier correction.</strong></p>\n<p>In principle, it sounds straightforward to make a simple projection or projections of where a point should be predicted to lie, and compare with the actual prediction. There’s a good notebook<br>\n<a href=\"https://www.kaggle.com/code/ravishah1/gsdc2-savgol-filter-outlier-removal-with-bo\" target=\"_blank\">GSDC2 - Savgol Filter + Outlier Removal with BO</a><br>\nthat implements an algorithm of this kind. However, I found no improvements from applying this kind of approach to what I’d consider “reasonably good” submissions. Even when I focused on fixing just the most egregious 100 or so outliers, I still made my score worse. I suspect it is hard to do what the competition metric requires, which is essentially to fix a “worst 5%” prediction well enough to move it out of the highest error 95%. I had some further ideas as to how to get this to work, but with the likely payoff rather small, I did not pursue them.</p>\n<p>Public score was not improved. One of my outlier correction models would marginally have improved the private score.</p>",
      "rawMarkdown": "Thanks to the hosts and to eveyone who took part and contributed to this very interesting and challenging competition. I am very much looking forward to reading through write-ups from the gold medal teams. As you will gather from what follows, I concentrated on firstly getting RTKLIB to work for me, and secondly on adding layers of postprocessing.\n\n**Step 0.  My starting point was a slightly adjusted version of @taroz1461’s notebook**\n\n[Carrier Smoothing + Robust WLS + Kalman Smoother](https://www.kaggle.com/code/taroz1461/carrier-smoothing-robust-wls-kalman-smoother)\n\nPublic score 2.964, Private score 3.462\n\nNote that I clearly had overfitted this to the public LB, as my version was worse than the original on the private set.\n\nSubsequent steps are listed in logical rather than temporal order.\n\n**Step 1. RTKLIB.**\n\n@timeverett provided a notebook\n\n[Getting Started with RTKLIB](https://www.kaggle.com/code/timeverett/getting-started-with-rtklib)\n\n giving detailed instructions on the installation of RKTLIB. From the quoted LB scores, this offered an improvement of close to 0.3 m on the public LB and more than 1.0 m on the private set,  more than the cumulative effect of the other steps of my workflow. Nonetheless, it was far from straightforward to get the software working correctly in my Windows environment. Most of my issues arose from the code not liking Windows paths. My initial problem turned out to be that I'd tried to run the code from a Windows folder with a space in its name, the space arising from the formatting of my user name, which threw an error. I therefore created a second instance, this time in a folder with no space in the name. The python code needed to be told explicitly about Windows paths. Thus, in step 4 of the process defined in @timeverett's notebook (so step 4 of step 1 in my nomenclature - sorry!), I changed the python code to explicitly add the required paths in the correct OS format. After three separate days of struggle, I finally got the code to run correctly.\n\nPublic score 2.682, Private score 2.189\n\n**Step 2. Generic Bias Correction**\n\nThe excellent notebooks\n\n[GSDC - Bias EDA](https://www.kaggle.com/code/saitodevel01/gsdc-bias-eda)\n\nand\n\n[GSDC - Bias Correction](https://www.kaggle.com/code/saitodevel01/gsdc-bias-correction)\n\nby @saitodevel01, written for the previous GSDC competition in 2021, demonstrate how the positions of the phone and driver in the car create biases between the recorded and actual positions in both the front-back and left-right directions.\nIn outline, the phone has better sightlines to satellites located in front of the car, since it is placed close to the to front windscreen, whereas sightlines of signals from satellites located behind the car may be compromised. Similarly, the driver sits to the left of the phone (these data came from the USA) and may block signals from satellites located to the left, whereas those to the right should have better communication. This would be expected to lead to an artefactual error locating the phone too far forward and too far to the right; hence the anticipated correction would move the phone backwards and leftwards.\n\nI analysed @saitodevel01's work from the 2021 competition, here is an example plot showing @saitodevel01's latitude correction against the cosine of the angle between the car's current direction and due north:\n\n<img src=\"https://i.imgur.com/HjF7Nmh.png\" width=\"400\" />\n\nThe corresponding plot for longitude looks broadly similar in form. My simplification is to treat the ellipse as the signal I want to model, and ignore all other points. That leads to a constant \"lever arm\" type correction of these biases, which I express as simple trigonometric functions of the direction the car is pointing in any one time step. Thus, I can compute latitude and longitude corrections for each predicted position.\n\nMy analysis of the prior work suggests that the predicted uniform generic correction is:\n\n40 cm from front to back; 27 cm from right to left.\n\nI added these corrections to the originally predicted latitude and longitude to generate a trial postprocessed submission.\nHowever, some trial and error showed that the best public LB scores came from a somewhat larger front-back correction of about 1.3 times that previously suggested. This works out as about 52 cm front to back correction.\n\nFor the left-right correction, surprisingly the best empirical public LB results came from a small correction of about 8 cm from left to right, that is in the counterintuitive direction. However, for simplicity, I set this generic left-right correction to zero.\n\nPublic score 2.595, Private score 2.158\n\n\n**Step 3: Phone-Specific Bias Correction**\n\nSome consideration of possible relevant effects led me to consider whether the bias correction might depend on the model of phone. There are various possible explanations for this. Possibly some aspect of the phone's hardware or software, or alternatively some quirk of the individual devices used in the experiment, meant that some phones were better able than others to communicate with satellites via compromised sightlines or through atmospheric or environmental interference. Hence, corrections could vary between phones. Alternatively, if each phone has a designated holder or position in the car, its results may be affected by sightlines and/or lever arm correction errors related to that position. For example. one phone holder may have better views backwards or to the left than another, depending on the extent to which the driver, seats, or experimental equipment block the view.  This seemed worth investigating, so a brief analysis was made using training set data, giving these results:\n\nGooglePixel4    \nForward error   -0.1798\nRelative to dataset -0.1639\nLeft error  -0.3383\nRelative to dataset -0.0136\n    \n    \nGooglePixel4XL  \nForward error   0.1375\nRelative to dataset 0.1535\nLeft error  -0.2301\nRelative to dataset 0.0947\n    \n    \nGooglePixel5    \nForward error   -0.7068\nRelative to dataset -0.6909\nLeft error  0.2203\nRelative to dataset 0.5450\n    \n    \nSamsungGalaxyS20Ultra   \nForward error   0.1469\nRelative to dataset 0.1628\nLeft error  -1.0847\nRelative to dataset -0.7599\n    \n    \nXiaomiMi8   \nForward error   0.4559\nRelative to dataset 0.4718\nLeft error  -1.0740\nRelative to dataset -0.7493\n\n\nIt seemed that it was worth trying a phone-specific set of corrections. Firstly, I looked at the front-back corrections. I only accept these corrections if they are in the same direction as the training results, and (for the front-back correction) no larger than the original generic correction. We can do this for the four phone models that are common to the training and test data; GooglePixel4, GooglePixel5, XiaomiMi8 and SamsungGalaxyS20Ultra. There is some trial and error involved here; the two (direction and size) constraints mentioned above are designed to limit public-score chasing. I also implemented phone-specific left-right corrections.\n\nThis is implemented in the notebook\n\n[Phone-specific Bias Corrections (Clipped)](https://www.kaggle.com/code/jbomitchell/phone-specific-bias-corrections-clipped)\n\nPublic score 2.508, Private score 2.094\n\n\n**Step 4: Stop averaging.**\n\nAny car journey is likely to involve periods when the vehicle is not going anywhere, whether parked up at the beginning or end, stopped at a light, or waiting for a junction or road to become clear. When we look at the data, we see substantial sequences of points where the position only changes slightly, these drifts of a few cm generally being artefacts. Despite the car being stationary, the GNSS position still appears to wander around a little, which leads to the idea of averaging coordinates over sets of points where the car was stopped. A threshold of 1 m/s seemed to work fairly well for many of my predictions, with most submission scores being improved. Some, but not all, submissions benefitted from pasting the output of the stop averaging process directly back into another iteration of the same computation. However, the RTKLIB solutions could not be improved by stop averaging. \n\nPublic and private scores were not improved for my best submissions.\n\n\n**Step 5: Tectonic correction.**\n\nI implemented corrections for the change in location of the base stations due to tectonic movement. The relevant base stations were: SLAC for Bay Area trips, with ITRF2014 velocity:\nnorthward = 0.0164 m/yr, eastward = -0.0343 m/yr, upward = 0.0009 m/yr\nStation VDCY for Los Angeles trips, with ITRF 2014 velocity:\nnorthward = = 0.0131 m/yr, eastward = -0.0374 m/yr, upward = 0.0009 m/yr\nThe RTKLIB calculations underlying my submission used a single set of base station positions computed to be correct midway through the time period over which the data were being collected, VDCY being selected for Los Angeles trips and SLAC for Bay Area trips. The distance through which I need to shift the predication is proportional to the (signed) number of days before or after the reference date that the run takes place (N), with the annual position shift being multiplied by (N/365.25) and assigned the appropriate sign depending on the direction of the time shift.\nI didn’t end up choosing a submission including this correction, based on the public score, but in fact it would have given a small improvement on the private set. The notebook is:\n\n[Tectonic Correction](https://www.kaggle.com/code/jbomitchell/tectonic-correction)\n\nPublic score 2.508, Private score 2.093\n\n\n**Step 6. Outlier correction.**\n\nIn principle, it sounds straightforward to make a simple projection or projections of where a point should be predicted to lie, and compare with the actual prediction. There’s a good notebook\n[GSDC2 - Savgol Filter + Outlier Removal with BO](https://www.kaggle.com/code/ravishah1/gsdc2-savgol-filter-outlier-removal-with-bo)\nthat implements an algorithm of this kind. However, I found no improvements from applying this kind of approach to what I’d consider “reasonably good” submissions. Even when I focused on fixing just the most egregious 100 or so outliers, I still made my score worse. I suspect it is hard to do what the competition metric requires, which is essentially to fix a “worst 5%” prediction well enough to move it out of the highest error 95%. I had some further ideas as to how to get this to work, but with the likely payoff rather small, I did not pursue them.\n\nPublic score was not improved. One of my outlier correction models would marginally have improved the private score.",
      "votes": null
    },
    {
      "id": "1908659",
      "postDate": "08/21/2022 21:29:08",
      "content": "<p>Thank you John for sharing your write-up. Congratulations for your position on Smartphone Decimeter  Challenge.</p>",
      "rawMarkdown": "Thank you John for sharing your write-up. Congratulations for your position on Smartphone Decimeter  Challenge.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1908659,
      "author_name": "mpwolke",
      "author_url": "",
      "post_date": "08/21/2022 21:29:08",
      "content": "<p>Thank you John for sharing your write-up. Congratulations for your position on Smartphone Decimeter  Challenge.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1877095": "Thanks to the hosts and to eveyone who took part and contributed to this very interesting and challenging competition. I am very much looking forward to reading through write-ups from the gold medal teams. As you will gather from what follows, I concentrated on firstly getting RTKLIB to work for me, and secondly on adding layers of postprocessing.\n\n**Step 0.  My starting point was a slightly adjusted version of @taroz1461’s notebook**\n\n[Carrier Smoothing + Robust WLS + Kalman Smoother](https://www.kaggle.com/code/taroz1461/carrier-smoothing-robust-wls-kalman-smoother)\n\nPublic score 2.964, Private score 3.462\n\nNote that I clearly had overfitted this to the public LB, as my version was worse than the original on the private set.\n\nSubsequent steps are listed in logical rather than temporal order.\n\n**Step 1. RTKLIB.**\n\n@timeverett provided a notebook\n\n[Getting Started with RTKLIB](https://www.kaggle.com/code/timeverett/getting-started-with-rtklib)\n\n giving detailed instructions on the installation of RKTLIB. From the quoted LB scores, this offered an improvement of close to 0.3 m on the public LB and more than 1.0 m on the private set,  more than the cumulative effect of the other steps of my workflow. Nonetheless, it was far from straightforward to get the software working correctly in my Windows environment. Most of my issues arose from the code not liking Windows paths. My initial problem turned out to be that I'd tried to run the code from a Windows folder with a space in its name, the space arising from the formatting of my user name, which threw an error. I therefore created a second instance, this time in a folder with no space in the name. The python code needed to be told explicitly about Windows paths. Thus, in step 4 of the process defined in @timeverett's notebook (so step 4 of step 1 in my nomenclature - sorry!), I changed the python code to explicitly add the required paths in the correct OS format. After three separate days of struggle, I finally got the code to run correctly.\n\nPublic score 2.682, Private score 2.189\n\n**Step 2. Generic Bias Correction**\n\nThe excellent notebooks\n\n[GSDC - Bias EDA](https://www.kaggle.com/code/saitodevel01/gsdc-bias-eda)\n\nand\n\n[GSDC - Bias Correction](https://www.kaggle.com/code/saitodevel01/gsdc-bias-correction)\n\nby @saitodevel01, written for the previous GSDC competition in 2021, demonstrate how the positions of the phone and driver in the car create biases between the recorded and actual positions in both the front-back and left-right directions.\nIn outline, the phone has better sightlines to satellites located in front of the car, since it is placed close to the to front windscreen, whereas sightlines of signals from satellites located behind the car may be compromised. Similarly, the driver sits to the left of the phone (these data came from the USA) and may block signals from satellites located to the left, whereas those to the right should have better communication. This would be expected to lead to an artefactual error locating the phone too far forward and too far to the right; hence the anticipated correction would move the phone backwards and leftwards.\n\nI analysed @saitodevel01's work from the 2021 competition, here is an example plot showing @saitodevel01's latitude correction against the cosine of the angle between the car's current direction and due north:\n\n<img src=\"https://i.imgur.com/HjF7Nmh.png\" width=\"400\" />\n\nThe corresponding plot for longitude looks broadly similar in form. My simplification is to treat the ellipse as the signal I want to model, and ignore all other points. That leads to a constant \"lever arm\" type correction of these biases, which I express as simple trigonometric functions of the direction the car is pointing in any one time step. Thus, I can compute latitude and longitude corrections for each predicted position.\n\nMy analysis of the prior work suggests that the predicted uniform generic correction is:\n\n40 cm from front to back; 27 cm from right to left.\n\nI added these corrections to the originally predicted latitude and longitude to generate a trial postprocessed submission.\nHowever, some trial and error showed that the best public LB scores came from a somewhat larger front-back correction of about 1.3 times that previously suggested. This works out as about 52 cm front to back correction.\n\nFor the left-right correction, surprisingly the best empirical public LB results came from a small correction of about 8 cm from left to right, that is in the counterintuitive direction. However, for simplicity, I set this generic left-right correction to zero.\n\nPublic score 2.595, Private score 2.158\n\n\n**Step 3: Phone-Specific Bias Correction**\n\nSome consideration of possible relevant effects led me to consider whether the bias correction might depend on the model of phone. There are various possible explanations for this. Possibly some aspect of the phone's hardware or software, or alternatively some quirk of the individual devices used in the experiment, meant that some phones were better able than others to communicate with satellites via compromised sightlines or through atmospheric or environmental interference. Hence, corrections could vary between phones. Alternatively, if each phone has a designated holder or position in the car, its results may be affected by sightlines and/or lever arm correction errors related to that position. For example. one phone holder may have better views backwards or to the left than another, depending on the extent to which the driver, seats, or experimental equipment block the view.  This seemed worth investigating, so a brief analysis was made using training set data, giving these results:\n\nGooglePixel4    \nForward error   -0.1798\nRelative to dataset -0.1639\nLeft error  -0.3383\nRelative to dataset -0.0136\n    \n    \nGooglePixel4XL  \nForward error   0.1375\nRelative to dataset 0.1535\nLeft error  -0.2301\nRelative to dataset 0.0947\n    \n    \nGooglePixel5    \nForward error   -0.7068\nRelative to dataset -0.6909\nLeft error  0.2203\nRelative to dataset 0.5450\n    \n    \nSamsungGalaxyS20Ultra   \nForward error   0.1469\nRelative to dataset 0.1628\nLeft error  -1.0847\nRelative to dataset -0.7599\n    \n    \nXiaomiMi8   \nForward error   0.4559\nRelative to dataset 0.4718\nLeft error  -1.0740\nRelative to dataset -0.7493\n\n\nIt seemed that it was worth trying a phone-specific set of corrections. Firstly, I looked at the front-back corrections. I only accept these corrections if they are in the same direction as the training results, and (for the front-back correction) no larger than the original generic correction. We can do this for the four phone models that are common to the training and test data; GooglePixel4, GooglePixel5, XiaomiMi8 and SamsungGalaxyS20Ultra. There is some trial and error involved here; the two (direction and size) constraints mentioned above are designed to limit public-score chasing. I also implemented phone-specific left-right corrections.\n\nThis is implemented in the notebook\n\n[Phone-specific Bias Corrections (Clipped)](https://www.kaggle.com/code/jbomitchell/phone-specific-bias-corrections-clipped)\n\nPublic score 2.508, Private score 2.094\n\n\n**Step 4: Stop averaging.**\n\nAny car journey is likely to involve periods when the vehicle is not going anywhere, whether parked up at the beginning or end, stopped at a light, or waiting for a junction or road to become clear. When we look at the data, we see substantial sequences of points where the position only changes slightly, these drifts of a few cm generally being artefacts. Despite the car being stationary, the GNSS position still appears to wander around a little, which leads to the idea of averaging coordinates over sets of points where the car was stopped. A threshold of 1 m/s seemed to work fairly well for many of my predictions, with most submission scores being improved. Some, but not all, submissions benefitted from pasting the output of the stop averaging process directly back into another iteration of the same computation. However, the RTKLIB solutions could not be improved by stop averaging. \n\nPublic and private scores were not improved for my best submissions.\n\n\n**Step 5: Tectonic correction.**\n\nI implemented corrections for the change in location of the base stations due to tectonic movement. The relevant base stations were: SLAC for Bay Area trips, with ITRF2014 velocity:\nnorthward = 0.0164 m/yr, eastward = -0.0343 m/yr, upward = 0.0009 m/yr\nStation VDCY for Los Angeles trips, with ITRF 2014 velocity:\nnorthward = = 0.0131 m/yr, eastward = -0.0374 m/yr, upward = 0.0009 m/yr\nThe RTKLIB calculations underlying my submission used a single set of base station positions computed to be correct midway through the time period over which the data were being collected, VDCY being selected for Los Angeles trips and SLAC for Bay Area trips. The distance through which I need to shift the predication is proportional to the (signed) number of days before or after the reference date that the run takes place (N), with the annual position shift being multiplied by (N/365.25) and assigned the appropriate sign depending on the direction of the time shift.\nI didn’t end up choosing a submission including this correction, based on the public score, but in fact it would have given a small improvement on the private set. The notebook is:\n\n[Tectonic Correction](https://www.kaggle.com/code/jbomitchell/tectonic-correction)\n\nPublic score 2.508, Private score 2.093\n\n\n**Step 6. Outlier correction.**\n\nIn principle, it sounds straightforward to make a simple projection or projections of where a point should be predicted to lie, and compare with the actual prediction. There’s a good notebook\n[GSDC2 - Savgol Filter + Outlier Removal with BO](https://www.kaggle.com/code/ravishah1/gsdc2-savgol-filter-outlier-removal-with-bo)\nthat implements an algorithm of this kind. However, I found no improvements from applying this kind of approach to what I’d consider “reasonably good” submissions. Even when I focused on fixing just the most egregious 100 or so outliers, I still made my score worse. I suspect it is hard to do what the competition metric requires, which is essentially to fix a “worst 5%” prediction well enough to move it out of the highest error 95%. I had some further ideas as to how to get this to work, but with the likely payoff rather small, I did not pursue them.\n\nPublic score was not improved. One of my outlier correction models would marginally have improved the private score.",
    "1908659": "Thank you John for sharing your write-up. Congratulations for your position on Smartphone Decimeter  Challenge."
  },
  "source": "meta"
}