{
  "id": 325702,
  "title": "📐 Deriving baseline WLS solutions",
  "url": "/competitions/smartphone-decimeter-2022/discussion/325702",
  "author_name": "",
  "post_date": "2022-05-17T21:53:58.867301600Z",
  "votes": 7,
  "comment_count": 4,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> has <a href=\"https://www.kaggle.com/code/junkoda/deriving-baseline-wls-positions-in-progress/notebook\" target=\"_blank\">a great notebook</a> that's trying to understand &amp; reproduce how Google has calculated the baseline ECEF positions that are provided to us.  This seems important for really getting to grips with the data.  Therefore, I'd like to follow up on that from 2 angles.</p>\n<p><strong>Firstly</strong>: <a href=\"https://www.kaggle.com/mohammedkhider\" target=\"_blank\">@mohammedkhider</a> / <a href=\"https://www.kaggle.com/gymf123\" target=\"_blank\">@gymf123</a> - Please can you provide a code sample that reproduces the baseline from the raw data?  Failing that, please can you provide a detailed description (down to the raw maths level) of how the baseline is created from the raw data?</p>\n<p><strong>Secondly</strong>: <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> - I've seen (on an Android page somewhere that I sadly now can't find again) that there's at least 1 WLS solver that excludes (a) unhealthy satellites and (b) satellites that are less than 5 degrees above the horizon.  I wonder if they've done the same here?  Perhaps that explains the difference between the baseline results and your notebook.</p>",
  "messages": [
    {
      "id": "1793427",
      "postDate": "05/17/2022 21:53:58",
      "content": "<p><a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> has <a href=\"https://www.kaggle.com/code/junkoda/deriving-baseline-wls-positions-in-progress/notebook\" target=\"_blank\">a great notebook</a> that's trying to understand &amp; reproduce how Google has calculated the baseline ECEF positions that are provided to us.  This seems important for really getting to grips with the data.  Therefore, I'd like to follow up on that from 2 angles.</p>\n<p><strong>Firstly</strong>: <a href=\"https://www.kaggle.com/mohammedkhider\" target=\"_blank\">@mohammedkhider</a> / <a href=\"https://www.kaggle.com/gymf123\" target=\"_blank\">@gymf123</a> - Please can you provide a code sample that reproduces the baseline from the raw data?  Failing that, please can you provide a detailed description (down to the raw maths level) of how the baseline is created from the raw data?</p>\n<p><strong>Secondly</strong>: <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> - I've seen (on an Android page somewhere that I sadly now can't find again) that there's at least 1 WLS solver that excludes (a) unhealthy satellites and (b) satellites that are less than 5 degrees above the horizon.  I wonder if they've done the same here?  Perhaps that explains the difference between the baseline results and your notebook.</p>",
      "rawMarkdown": "junkoda has [a great notebook](https://www.kaggle.com/code/junkoda/deriving-baseline-wls-positions-in-progress/notebook) that's trying to understand & reproduce how Google has calculated the baseline ECEF positions that are provided to us.  This seems important for really getting to grips with the data.  Therefore, I'd like to follow up on that from 2 angles.\n\n**Firstly**: @mohammedkhider / @gymf123 - Please can you provide a code sample that reproduces the baseline from the raw data?  Failing that, please can you provide a detailed description (down to the raw maths level) of how the baseline is created from the raw data?\n\n**Secondly**: @junkoda - I've seen (on an Android page somewhere that I sadly now can't find again) that there's at least 1 WLS solver that excludes (a) unhealthy satellites and (b) satellites that are less than 5 degrees above the horizon.  I wonder if they've done the same here?  Perhaps that explains the difference between the baseline results and your notebook.",
      "votes": null
    },
    {
      "id": "1793447",
      "postDate": "05/17/2022 22:17:46",
      "content": "<p>Following up on the second point and looking at the specific example you use in your notebook…</p>\n<p>You can get the residuals (i.e. per-satellite pseudorange differences from the final position) from the WLS optimizer output (using <code>opt.fun</code>).  You'll see that most are &lt;1m off but there's one satellite (with <code>Svid</code> 21) that has a residual of &gt;7m.  Excluding this single satellite and re-running the optimizer brings the final error down to 0.35m.  I appreciate that still isn't 0, and my choice of excluding this satellite was arbitrary.  However, if there's a principled way to know which satellites they've excluded then perhaps we'll be able to reproduce exactly.</p>\n<hr>\n<p>Edit: Sure enough, satellite 21 is much lower in the sky.  The data we're provided says 6.8 degrees.  Sadly this doesn't quite match with elevation calculated (from the provided reference position) using <a href=\"https://gis.stackexchange.com/questions/58923/calculating-view-angle\" target=\"_blank\">this post</a> - which gives 6.6 degrees.  But that's above the 5 degree cut-off I read about.  Also, satellite 21 already has a high uncertainty (and therefore low weight) so perhaps this is all a red herring.  Shame that there's another calculation I can't reproduce though.</p>\n<p>I've tried re-running WLS after excluding all satellites below a certain elevation threshold and all satellites below an uncertainty threshold (trying different values of the thresholds).  However, I can't reproduce the provided baseline with either of these. </p>",
      "rawMarkdown": "Following up on the second point and looking at the specific example you use in your notebook...\n\nYou can get the residuals (i.e. per-satellite pseudorange differences from the final position) from the WLS optimizer output (using `opt.fun`).  You'll see that most are <1m off but there's one satellite (with `Svid` 21) that has a residual of >7m.  Excluding this single satellite and re-running the optimizer brings the final error down to 0.35m.  I appreciate that still isn't 0, and my choice of excluding this satellite was arbitrary.  However, if there's a principled way to know which satellites they've excluded then perhaps we'll be able to reproduce exactly.\n\n---\n\nEdit: Sure enough, satellite 21 is much lower in the sky.  The data we're provided says 6.8 degrees.  Sadly this doesn't quite match with elevation calculated (from the provided reference position) using [this post](https://gis.stackexchange.com/questions/58923/calculating-view-angle) - which gives 6.6 degrees.  But that's above the 5 degree cut-off I read about.  Also, satellite 21 already has a high uncertainty (and therefore low weight) so perhaps this is all a red herring.  Shame that there's another calculation I can't reproduce though.\n\nI've tried re-running WLS after excluding all satellites below a certain elevation threshold and all satellites below an uncertainty threshold (trying different values of the thresholds).  However, I can't reproduce the provided baseline with either of these.",
      "votes": null
    },
    {
      "id": "1794534",
      "postDate": "05/18/2022 22:53:59",
      "content": "<p>Thanks for mentioning my notebook, working on understanding WLS, and checking the satellite elevation hypothesis!</p>\n<p>I realized that many criteria for invalid measurements in <a href=\"https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/238583\" target=\"_blank\">last year's WLS method</a> are the same as the condition for supplemental/rinex.o in the data section. E.g., some data in gnss have low signal to noise: \"CN0 is less than 20 dB-Hz\". I tried to select the same data as rinex.o by finding missing pseudo ranges in this file, but haven't improved the match in position so far. (Also, there is no \"invalid data\" in the first epoch of the first drive.)</p>\n<p>Selection based on <code>CarrierFrequencyHz</code> improves the match, but not by orders of magnitude <br>\nfor all epochs in the first drive.</p>",
      "rawMarkdown": "Thanks for mentioning my notebook, working on understanding WLS, and checking the satellite elevation hypothesis!\n\nI realized that many criteria for invalid measurements in [last year's WLS method](https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/238583) are the same as the condition for supplemental/rinex.o in the data section. E.g., some data in gnss have low signal to noise: \"CN0 is less than 20 dB-Hz\". I tried to select the same data as rinex.o by finding missing pseudo ranges in this file, but haven't improved the match in position so far. (Also, there is no \"invalid data\" in the first epoch of the first drive.)\n\nSelection based on `CarrierFrequencyHz` improves the match, but not by orders of magnitude \nfor all epochs in the first drive.",
      "votes": null
    },
    {
      "id": "1803936",
      "postDate": "05/28/2022 10:56:45",
      "content": "<p>maybe these are useful:</p>\n<p><a href=\"https://webthesis.biblio.polito.it/11702/1/tesi.pdf\" target=\"_blank\">https://webthesis.biblio.polito.it/11702/1/tesi.pdf</a></p>\n<p>section 2.2 and section 4 tries to explain google code:<br>\n<a href=\"https://github.com/google/gps-measurement-tools/blob/master/GNSSLogger/pseudorange/src/main/java/com/google/location/lbs/gnss/gps/pseudorange/UserPositionVelocityWeightedLeastSquare.java\" target=\"_blank\">https://github.com/google/gps-measurement-tools/blob/master/GNSSLogger/pseudorange/src/main/java/com/google/location/lbs/gnss/gps/pseudorange/UserPositionVelocityWeightedLeastSquare.java</a></p>\n<p>related:<br>\n<a href=\"https://insidegnss.com/gnss-analysis-tools-from-google/\" target=\"_blank\">https://insidegnss.com/gnss-analysis-tools-from-google/</a></p>",
      "rawMarkdown": "maybe these are useful:\n\nhttps://webthesis.biblio.polito.it/11702/1/tesi.pdf\n\nsection 2.2 and section 4 tries to explain google code:\nhttps://github.com/google/gps-measurement-tools/blob/master/GNSSLogger/pseudorange/src/main/java/com/google/location/lbs/gnss/gps/pseudorange/UserPositionVelocityWeightedLeastSquare.java\n\n\nrelated:\nhttps://insidegnss.com/gnss-analysis-tools-from-google/",
      "votes": null
    },
    {
      "id": "1805363",
      "postDate": "05/30/2022 05:05:10",
      "content": "<p><strong>RawPseudorangeMeters all NaN</strong></p>\n<p>There are only 33 NaNs in the entire <code>WlsPositionXEcefMeters</code> in the training set (7 in test), but<br>\nthere are a lot more epochs that all <code>RawPseudorangeMeters</code> in an epoch (timestep) are NaN. In such cases, I confirmed that <code>WlsPositionXEcefMeters</code> are filled by previous finite values.</p>\n<p>Epochs with NaN-only pseudo ranges / total number of epochs</p>\n<pre><code>787 / 295633 (train)\n104 / 66097 (test)\n</code></pre>\n<p>There are a lot of NaNs in RawPseudorangeMeters, but epochs with all NaN pseudo ranges are minor and not a serious problem unless you submit NaN as a prediction.</p>\n<p>The filled values are not exactly equal but differ by machine precision ~ 1e-10, which indicates that they are not substituted but recalculated with identical data in a non-deterministic way (e.g., multi-core parallel)</p>",
      "rawMarkdown": "**RawPseudorangeMeters all NaN**\n\nThere are only 33 NaNs in the entire `WlsPositionXEcefMeters` in the training set (7 in test), but\nthere are a lot more epochs that all `RawPseudorangeMeters` in an epoch (timestep) are NaN. In such cases, I confirmed that `WlsPositionXEcefMeters` are filled by previous finite values.\n\nEpochs with NaN-only pseudo ranges / total number of epochs\n\n```\n787 / 295633 (train)\n104 / 66097 (test)\n```\n\nThere are a lot of NaNs in RawPseudorangeMeters, but epochs with all NaN pseudo ranges are minor and not a serious problem unless you submit NaN as a prediction.\n\nThe filled values are not exactly equal but differ by machine precision ~ 1e-10, which indicates that they are not substituted but recalculated with identical data in a non-deterministic way (e.g., multi-core parallel)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1793447,
      "author_name": "andrewrrose",
      "author_url": "",
      "post_date": "05/17/2022 22:17:46",
      "content": "<p>Following up on the second point and looking at the specific example you use in your notebook…</p>\n<p>You can get the residuals (i.e. per-satellite pseudorange differences from the final position) from the WLS optimizer output (using <code>opt.fun</code>).  You'll see that most are &lt;1m off but there's one satellite (with <code>Svid</code> 21) that has a residual of &gt;7m.  Excluding this single satellite and re-running the optimizer brings the final error down to 0.35m.  I appreciate that still isn't 0, and my choice of excluding this satellite was arbitrary.  However, if there's a principled way to know which satellites they've excluded then perhaps we'll be able to reproduce exactly.</p>\n<hr>\n<p>Edit: Sure enough, satellite 21 is much lower in the sky.  The data we're provided says 6.8 degrees.  Sadly this doesn't quite match with elevation calculated (from the provided reference position) using <a href=\"https://gis.stackexchange.com/questions/58923/calculating-view-angle\" target=\"_blank\">this post</a> - which gives 6.6 degrees.  But that's above the 5 degree cut-off I read about.  Also, satellite 21 already has a high uncertainty (and therefore low weight) so perhaps this is all a red herring.  Shame that there's another calculation I can't reproduce though.</p>\n<p>I've tried re-running WLS after excluding all satellites below a certain elevation threshold and all satellites below an uncertainty threshold (trying different values of the thresholds).  However, I can't reproduce the provided baseline with either of these. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1794534,
      "author_name": "junkoda",
      "author_url": "",
      "post_date": "05/18/2022 22:53:59",
      "content": "<p>Thanks for mentioning my notebook, working on understanding WLS, and checking the satellite elevation hypothesis!</p>\n<p>I realized that many criteria for invalid measurements in <a href=\"https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/238583\" target=\"_blank\">last year's WLS method</a> are the same as the condition for supplemental/rinex.o in the data section. E.g., some data in gnss have low signal to noise: \"CN0 is less than 20 dB-Hz\". I tried to select the same data as rinex.o by finding missing pseudo ranges in this file, but haven't improved the match in position so far. (Also, there is no \"invalid data\" in the first epoch of the first drive.)</p>\n<p>Selection based on <code>CarrierFrequencyHz</code> improves the match, but not by orders of magnitude <br>\nfor all epochs in the first drive.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1803936,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/28/2022 10:56:45",
      "content": "<p>maybe these are useful:</p>\n<p><a href=\"https://webthesis.biblio.polito.it/11702/1/tesi.pdf\" target=\"_blank\">https://webthesis.biblio.polito.it/11702/1/tesi.pdf</a></p>\n<p>section 2.2 and section 4 tries to explain google code:<br>\n<a href=\"https://github.com/google/gps-measurement-tools/blob/master/GNSSLogger/pseudorange/src/main/java/com/google/location/lbs/gnss/gps/pseudorange/UserPositionVelocityWeightedLeastSquare.java\" target=\"_blank\">https://github.com/google/gps-measurement-tools/blob/master/GNSSLogger/pseudorange/src/main/java/com/google/location/lbs/gnss/gps/pseudorange/UserPositionVelocityWeightedLeastSquare.java</a></p>\n<p>related:<br>\n<a href=\"https://insidegnss.com/gnss-analysis-tools-from-google/\" target=\"_blank\">https://insidegnss.com/gnss-analysis-tools-from-google/</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1805363,
      "author_name": "junkoda",
      "author_url": "",
      "post_date": "05/30/2022 05:05:10",
      "content": "<p><strong>RawPseudorangeMeters all NaN</strong></p>\n<p>There are only 33 NaNs in the entire <code>WlsPositionXEcefMeters</code> in the training set (7 in test), but<br>\nthere are a lot more epochs that all <code>RawPseudorangeMeters</code> in an epoch (timestep) are NaN. In such cases, I confirmed that <code>WlsPositionXEcefMeters</code> are filled by previous finite values.</p>\n<p>Epochs with NaN-only pseudo ranges / total number of epochs</p>\n<pre><code>787 / 295633 (train)\n104 / 66097 (test)\n</code></pre>\n<p>There are a lot of NaNs in RawPseudorangeMeters, but epochs with all NaN pseudo ranges are minor and not a serious problem unless you submit NaN as a prediction.</p>\n<p>The filled values are not exactly equal but differ by machine precision ~ 1e-10, which indicates that they are not substituted but recalculated with identical data in a non-deterministic way (e.g., multi-core parallel)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1793427": "junkoda has [a great notebook](https://www.kaggle.com/code/junkoda/deriving-baseline-wls-positions-in-progress/notebook) that's trying to understand & reproduce how Google has calculated the baseline ECEF positions that are provided to us.  This seems important for really getting to grips with the data.  Therefore, I'd like to follow up on that from 2 angles.\n\n**Firstly**: @mohammedkhider / @gymf123 - Please can you provide a code sample that reproduces the baseline from the raw data?  Failing that, please can you provide a detailed description (down to the raw maths level) of how the baseline is created from the raw data?\n\n**Secondly**: @junkoda - I've seen (on an Android page somewhere that I sadly now can't find again) that there's at least 1 WLS solver that excludes (a) unhealthy satellites and (b) satellites that are less than 5 degrees above the horizon.  I wonder if they've done the same here?  Perhaps that explains the difference between the baseline results and your notebook.",
    "1793447": "Following up on the second point and looking at the specific example you use in your notebook...\n\nYou can get the residuals (i.e. per-satellite pseudorange differences from the final position) from the WLS optimizer output (using `opt.fun`).  You'll see that most are <1m off but there's one satellite (with `Svid` 21) that has a residual of >7m.  Excluding this single satellite and re-running the optimizer brings the final error down to 0.35m.  I appreciate that still isn't 0, and my choice of excluding this satellite was arbitrary.  However, if there's a principled way to know which satellites they've excluded then perhaps we'll be able to reproduce exactly.\n\n---\n\nEdit: Sure enough, satellite 21 is much lower in the sky.  The data we're provided says 6.8 degrees.  Sadly this doesn't quite match with elevation calculated (from the provided reference position) using [this post](https://gis.stackexchange.com/questions/58923/calculating-view-angle) - which gives 6.6 degrees.  But that's above the 5 degree cut-off I read about.  Also, satellite 21 already has a high uncertainty (and therefore low weight) so perhaps this is all a red herring.  Shame that there's another calculation I can't reproduce though.\n\nI've tried re-running WLS after excluding all satellites below a certain elevation threshold and all satellites below an uncertainty threshold (trying different values of the thresholds).  However, I can't reproduce the provided baseline with either of these.",
    "1794534": "Thanks for mentioning my notebook, working on understanding WLS, and checking the satellite elevation hypothesis!\n\nI realized that many criteria for invalid measurements in [last year's WLS method](https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/238583) are the same as the condition for supplemental/rinex.o in the data section. E.g., some data in gnss have low signal to noise: \"CN0 is less than 20 dB-Hz\". I tried to select the same data as rinex.o by finding missing pseudo ranges in this file, but haven't improved the match in position so far. (Also, there is no \"invalid data\" in the first epoch of the first drive.)\n\nSelection based on `CarrierFrequencyHz` improves the match, but not by orders of magnitude \nfor all epochs in the first drive.",
    "1803936": "maybe these are useful:\n\nhttps://webthesis.biblio.polito.it/11702/1/tesi.pdf\n\nsection 2.2 and section 4 tries to explain google code:\nhttps://github.com/google/gps-measurement-tools/blob/master/GNSSLogger/pseudorange/src/main/java/com/google/location/lbs/gnss/gps/pseudorange/UserPositionVelocityWeightedLeastSquare.java\n\n\nrelated:\nhttps://insidegnss.com/gnss-analysis-tools-from-google/",
    "1805363": "**RawPseudorangeMeters all NaN**\n\nThere are only 33 NaNs in the entire `WlsPositionXEcefMeters` in the training set (7 in test), but\nthere are a lot more epochs that all `RawPseudorangeMeters` in an epoch (timestep) are NaN. In such cases, I confirmed that `WlsPositionXEcefMeters` are filled by previous finite values.\n\nEpochs with NaN-only pseudo ranges / total number of epochs\n\n```\n787 / 295633 (train)\n104 / 66097 (test)\n```\n\nThere are a lot of NaNs in RawPseudorangeMeters, but epochs with all NaN pseudo ranges are minor and not a serious problem unless you submit NaN as a prediction.\n\nThe filled values are not exactly equal but differ by machine precision ~ 1e-10, which indicates that they are not substituted but recalculated with identical data in a non-deterministic way (e.g., multi-core parallel)"
  },
  "source": "meta"
}