{
  "id": 261904,
  "title": "6th Place Solution (shimacos part)",
  "url": "/competitions/google-smartphone-decimeter-challenge/writeups/lightspm-6th-place-solution-shimacos-part",
  "author_name": "",
  "post_date": "2021-08-06T02:17:28.133Z",
  "votes": 29,
  "comment_count": 2,
  "views": 0,
  "content": "<p>First of all, I would like to thank hosts for organizing such a unique competition!<br>\nAnd in this competition, I’m able to become a Grand Master and my teammate <a href=\"https://www.kaggle.com/fuumin621\" target=\"_blank\">@fuumin621</a> is able to become a master!<br>\nHere is an introduction to my part of our brief solution.<br>\nIt was important for us to use AccumulatedDeltaRange properly and look at the data carefully..</p>\n<h1>Preprocess</h1>\n<p>I predicted the anomaly (distance from ground truth) and speed from IMU data and baseline data by lightGBM.<br>\nThen, I replaced anomalies above the threshold with NULLs and linearly interpolated between them.</p>\n<h1>Kalman smoothing</h1>\n<p>I applied two pattern Linear Kaman Smoothing.<br>\nOne takes into account velocity, and the other takes into account acceleration (i.e. 4D and 6D).<br>\nAlso, since the interval of the observed values was not constant, I replaced delta_t in the state transition matrix for each update step.<br>\nI tried making the observation noise larger in downtown, predicting the yaw rate, and using a nonlinear Kalman filter, but it didn't work.</p>\n<h1>AccumulatedDeltaRange</h1>\n<p>I knew that phase observation was important, so I tried to improve the baseline using the AccumulatedDeltaRange from the satellite data, but it did not work due to lack of knowledge…<br>\nSo I decided using the <a href=\"https://developer.android.com/guide/topics/sensors/gnss\" target=\"_blank\">GNSS Analysis app from google</a>.<br>\nBy using this tool, I was able to obtain observations using the AccumulatedDeltaRange for some of the data.<br>\nThe observation data itself was difficult to use directly because it contained bias, but the relative distances obtained by taking the differences between them were very accurate like following image. (delta_adr = ADR(t) - ADR(t-1))<br>\n<img src=\"https://user-images.githubusercontent.com/24289602/128328750-1217fc5d-5280-42bd-be8b-0adf2485f3af.png\" alt=\"image\"><br>\nI used this to implement a forward hatch filter and a backward hatch filter and averaged them after applied Kaman smoothing.</p>\n<pre><code>   prev_smooth = lat_lngs[0]\n   res = [prev_smooth]\n   for raw, delta_adr, delta_t in zip(\n       lat_lngs[1:, :], delta_adrs[1:, :], delta_ts[1:]\n   ):\n       if not all(np.isnan(delta_adr)):\n           smooth = 1 / M * raw + (M - 1) / M * (\n               prev_smooth + delta_adr * delta_t\n           )\n       else:\n           smooth = raw\n       prev_smooth = smooth\n       res.append(prev_smooth)\n  res = np.array(res)\n</code></pre>\n<p>There is a parameter called M, so I optimized it using GroupKFold for each area. (Find the parameter that is the smallest in the train data and verify it with valid data.)<br>\nSmoothing with this hatch filter was very effective.<br>\nThere were many missing delta_adr values for some phones, so we filled them with values from other phones.</p>\n<h1>Weighted phone mean</h1>\n<p>After applied hatch filter, I weighted average latitude and longitude using the weights of phone optimized by GroupKFold as well as M.<br>\nSome of the <code>millisSinceGpsEpoch</code> had a little gap depending on the phone, so I used linear interpolation to fill it.</p>\n<h1>Snap2grid</h1>\n<p>In the downtown area, I performed a discrete optimization using linearly interpolated ground truth.<br>\nTo prevent overfitting, the ground truth in itself was not included in the target to be snapped.<br>\nThe objective function is as follows, and I used greedy search to find the optimal path.</p>\n<pre><code>cost = distance2cur + distance2prev * alpha\n</code></pre>\n<h1>Stay point mean</h1>\n<p>Using the predicted value of the speed of lightgbm, I took the average for points below the threshold.</p>\n<h1>Ensemble</h1>\n<p>Using various parameters, I took the average of them and the CV was about 2.149.<br>\nFinally, after adding other members' models and post-processing, we got CV  2.0569 and LB 2.285.<br>\nIn Private, I knew that there was a lot of weight on highways and trees, so I decided on the ensemble weights taking them into account.</p>",
  "messages": [
    {
      "id": "1451236",
      "postDate": "08/05/2021 09:42:52",
      "content": "<p>First of all, I would like to thank hosts for organizing such a unique competition!<br>\nAnd in this competition, I’m able to become a Grand Master and my teammate <a href=\"https://www.kaggle.com/fuumin621\" target=\"_blank\">@fuumin621</a> is able to become a master!<br>\nHere is an introduction to my part of our brief solution.<br>\nIt was important for us to use AccumulatedDeltaRange properly and look at the data carefully..</p>\n<h1>Preprocess</h1>\n<p>I predicted the anomaly (distance from ground truth) and speed from IMU data and baseline data by lightGBM.<br>\nThen, I replaced anomalies above the threshold with NULLs and linearly interpolated between them.</p>\n<h1>Kalman smoothing</h1>\n<p>I applied two pattern Linear Kaman Smoothing.<br>\nOne takes into account velocity, and the other takes into account acceleration (i.e. 4D and 6D).<br>\nAlso, since the interval of the observed values was not constant, I replaced delta_t in the state transition matrix for each update step.<br>\nI tried making the observation noise larger in downtown, predicting the yaw rate, and using a nonlinear Kalman filter, but it didn't work.</p>\n<h1>AccumulatedDeltaRange</h1>\n<p>I knew that phase observation was important, so I tried to improve the baseline using the AccumulatedDeltaRange from the satellite data, but it did not work due to lack of knowledge…<br>\nSo I decided using the <a href=\"https://developer.android.com/guide/topics/sensors/gnss\" target=\"_blank\">GNSS Analysis app from google</a>.<br>\nBy using this tool, I was able to obtain observations using the AccumulatedDeltaRange for some of the data.<br>\nThe observation data itself was difficult to use directly because it contained bias, but the relative distances obtained by taking the differences between them were very accurate like following image. (delta_adr = ADR(t) - ADR(t-1))<br>\n<img src=\"https://user-images.githubusercontent.com/24289602/128328750-1217fc5d-5280-42bd-be8b-0adf2485f3af.png\" alt=\"image\"><br>\nI used this to implement a forward hatch filter and a backward hatch filter and averaged them after applied Kaman smoothing.</p>\n<pre><code>   prev_smooth = lat_lngs[0]\n   res = [prev_smooth]\n   for raw, delta_adr, delta_t in zip(\n       lat_lngs[1:, :], delta_adrs[1:, :], delta_ts[1:]\n   ):\n       if not all(np.isnan(delta_adr)):\n           smooth = 1 / M * raw + (M - 1) / M * (\n               prev_smooth + delta_adr * delta_t\n           )\n       else:\n           smooth = raw\n       prev_smooth = smooth\n       res.append(prev_smooth)\n  res = np.array(res)\n</code></pre>\n<p>There is a parameter called M, so I optimized it using GroupKFold for each area. (Find the parameter that is the smallest in the train data and verify it with valid data.)<br>\nSmoothing with this hatch filter was very effective.<br>\nThere were many missing delta_adr values for some phones, so we filled them with values from other phones.</p>\n<h1>Weighted phone mean</h1>\n<p>After applied hatch filter, I weighted average latitude and longitude using the weights of phone optimized by GroupKFold as well as M.<br>\nSome of the <code>millisSinceGpsEpoch</code> had a little gap depending on the phone, so I used linear interpolation to fill it.</p>\n<h1>Snap2grid</h1>\n<p>In the downtown area, I performed a discrete optimization using linearly interpolated ground truth.<br>\nTo prevent overfitting, the ground truth in itself was not included in the target to be snapped.<br>\nThe objective function is as follows, and I used greedy search to find the optimal path.</p>\n<pre><code>cost = distance2cur + distance2prev * alpha\n</code></pre>\n<h1>Stay point mean</h1>\n<p>Using the predicted value of the speed of lightgbm, I took the average for points below the threshold.</p>\n<h1>Ensemble</h1>\n<p>Using various parameters, I took the average of them and the CV was about 2.149.<br>\nFinally, after adding other members' models and post-processing, we got CV  2.0569 and LB 2.285.<br>\nIn Private, I knew that there was a lot of weight on highways and trees, so I decided on the ensemble weights taking them into account.</p>",
      "rawMarkdown": "First of all, I would like to thank hosts for organizing such a unique competition!\nAnd in this competition, I’m able to become a Grand Master and my teammate @fuumin621 is able to become a master!\n\n\nHere is an introduction to my part of our brief solution.\nIt was important for us to use AccumulatedDeltaRange properly and look at the data carefully..\n\n# Preprocess\n\nI predicted the anomaly (distance from ground truth) and speed from IMU data and baseline data by lightGBM.\nThen, I replaced anomalies above the threshold with NULLs and linearly interpolated between them.\n\n# Kalman smoothing\n\nI applied two pattern Linear Kaman Smoothing.\nOne takes into account velocity, and the other takes into account acceleration (i.e. 4D and 6D).\nAlso, since the interval of the observed values was not constant, I replaced delta_t in the state transition matrix for each update step.\nI tried making the observation noise larger in downtown, predicting the yaw rate, and using a nonlinear Kalman filter, but it didn't work.\n\n\n# AccumulatedDeltaRange\n\nI knew that phase observation was important, so I tried to improve the baseline using the AccumulatedDeltaRange from the satellite data, but it did not work due to lack of knowledge…\nSo I decided using the [GNSS Analysis app from google] (https://developer.android.com/guide/topics/sensors/gnss).\nBy using this tool, I was able to obtain observations using the AccumulatedDeltaRange for some of the data.\nThe observation data itself was difficult to use directly because it contained bias, but the relative distances obtained by taking the differences between them were very accurate like following image. (delta_adr = ADR(t) - ADR(t-1))\n![image](https://user-images.githubusercontent.com/24289602/128328750-1217fc5d-5280-42bd-be8b-0adf2485f3af.png)\n\nI used this to implement a forward hatch filter and a backward hatch filter and averaged them after applied Kaman smoothing.\n```python\n    prev_smooth = lat_lngs[0]\n    res = [prev_smooth]\n    for raw, delta_adr, delta_t in zip(\n        lat_lngs[1:, :], delta_adrs[1:, :], delta_ts[1:]\n    ):\n        if not all(np.isnan(delta_adr)):\n            smooth = 1 / M * raw + (M - 1) / M * (\n                prev_smooth + delta_adr * delta_t\n            )\n        else:\n            smooth = raw\n        prev_smooth = smooth\n        res.append(prev_smooth)\n   res = np.array(res)\n```\nThere is a parameter called M, so I optimized it using GroupKFold for each area. (Find the parameter that is the smallest in the train data and verify it with valid data.)\nSmoothing with this hatch filter was very effective.\nThere were many missing delta_adr values for some phones, so we filled them with values from other phones.\n\n# Weighted phone mean\n\nAfter applied hatch filter, I weighted average latitude and longitude using the weights of phone optimized by GroupKFold as well as M.\nSome of the `millisSinceGpsEpoch` had a little gap depending on the phone, so I used linear interpolation to fill it.\n\n\n# Snap2grid\n\nIn the downtown area, I performed a discrete optimization using linearly interpolated ground truth.\nTo prevent overfitting, the ground truth in itself was not included in the target to be snapped.\nThe objective function is as follows, and I used greedy search to find the optimal path.\n```python\ncost = distance2cur + distance2prev * alpha\n```\n\n# Stay point mean\n\nUsing the predicted value of the speed of lightgbm, I took the average for points below the threshold.\n\n# Ensemble\n\nUsing various parameters, I took the average of them and the CV was about 2.149.\nFinally, after adding other members' models and post-processing, we got CV  2.0569 and LB 2.285.\nIn Private, I knew that there was a lot of weight on highways and trees, so I decided on the ensemble weights taking them into account.",
      "votes": null
    },
    {
      "id": "1451405",
      "postDate": "08/05/2021 10:36:12",
      "content": "<p>Good jod! i will learn from it!</p>",
      "rawMarkdown": "Good jod! i will learn from it!",
      "votes": null
    },
    {
      "id": "1454749",
      "postDate": "08/06/2021 10:04:33",
      "content": "<p>Congrats  GrandMaster, and It was fun to work with you.</p>\n<p>Our almost solution is written by Shimacos and <a href=\"https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/262195\" target=\"_blank\">monnu</a>, so I will add some other point related shimacos's solution. (especially ADR)</p>\n<h2>ADR shift</h2>\n<p>AccumulatedDeltaRange(ADR) is very strong about relative point relationship, but absolute point is a little weak.<br>\nTherefore I calculate the difference between ADR and baseline's average location (after postprocess) in collection / phone, and shift to fit baseline's average.<br>\nIt improved ADR's location performance.</p>\n<h2>Stop point correlation using ADR</h2>\n<p>ADR is also useful to stop point post processing.<br>\nUsing not only speed(predict by lightgbm) but also ADR change from previous point make final submission improve.</p>\n<ul>\n<li>When (speed &lt; threshold1) &amp; (adr_change &lt; threshold2), stop flag is true</li>\n<li>and there are more than N point in a row, these point replace to median point.</li>\n</ul>\n<p>This is done after ensemble, and improve model.</p>",
      "rawMarkdown": "Congrats  GrandMaster, and It was fun to work with you.\n\nOur almost solution is written by Shimacos and [monnu](https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/262195), so I will add some other point related shimacos's solution. (especially ADR)\n\n## ADR shift\n\nAccumulatedDeltaRange(ADR) is very strong about relative point relationship, but absolute point is a little weak.\nTherefore I calculate the difference between ADR and baseline's average location (after postprocess) in collection / phone, and shift to fit baseline's average.\nIt improved ADR's location performance.\n\n## Stop point correlation using ADR\n\nADR is also useful to stop point post processing.\nUsing not only speed(predict by lightgbm) but also ADR change from previous point make final submission improve.\n\n- When (speed < threshold1) & (adr_change < threshold2), stop flag is true\n- and there are more than N point in a row, these point replace to median point.\n\nThis is done after ensemble, and improve model.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1451405,
      "author_name": "guohey",
      "author_url": "",
      "post_date": "08/05/2021 10:36:12",
      "content": "<p>Good jod! i will learn from it!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1454749,
      "author_name": "go5kuramubon",
      "author_url": "",
      "post_date": "08/06/2021 10:04:33",
      "content": "<p>Congrats  GrandMaster, and It was fun to work with you.</p>\n<p>Our almost solution is written by Shimacos and <a href=\"https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/262195\" target=\"_blank\">monnu</a>, so I will add some other point related shimacos's solution. (especially ADR)</p>\n<h2>ADR shift</h2>\n<p>AccumulatedDeltaRange(ADR) is very strong about relative point relationship, but absolute point is a little weak.<br>\nTherefore I calculate the difference between ADR and baseline's average location (after postprocess) in collection / phone, and shift to fit baseline's average.<br>\nIt improved ADR's location performance.</p>\n<h2>Stop point correlation using ADR</h2>\n<p>ADR is also useful to stop point post processing.<br>\nUsing not only speed(predict by lightgbm) but also ADR change from previous point make final submission improve.</p>\n<ul>\n<li>When (speed &lt; threshold1) &amp; (adr_change &lt; threshold2), stop flag is true</li>\n<li>and there are more than N point in a row, these point replace to median point.</li>\n</ul>\n<p>This is done after ensemble, and improve model.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1451236": "First of all, I would like to thank hosts for organizing such a unique competition!\nAnd in this competition, I’m able to become a Grand Master and my teammate @fuumin621 is able to become a master!\n\n\nHere is an introduction to my part of our brief solution.\nIt was important for us to use AccumulatedDeltaRange properly and look at the data carefully..\n\n# Preprocess\n\nI predicted the anomaly (distance from ground truth) and speed from IMU data and baseline data by lightGBM.\nThen, I replaced anomalies above the threshold with NULLs and linearly interpolated between them.\n\n# Kalman smoothing\n\nI applied two pattern Linear Kaman Smoothing.\nOne takes into account velocity, and the other takes into account acceleration (i.e. 4D and 6D).\nAlso, since the interval of the observed values was not constant, I replaced delta_t in the state transition matrix for each update step.\nI tried making the observation noise larger in downtown, predicting the yaw rate, and using a nonlinear Kalman filter, but it didn't work.\n\n\n# AccumulatedDeltaRange\n\nI knew that phase observation was important, so I tried to improve the baseline using the AccumulatedDeltaRange from the satellite data, but it did not work due to lack of knowledge…\nSo I decided using the [GNSS Analysis app from google] (https://developer.android.com/guide/topics/sensors/gnss).\nBy using this tool, I was able to obtain observations using the AccumulatedDeltaRange for some of the data.\nThe observation data itself was difficult to use directly because it contained bias, but the relative distances obtained by taking the differences between them were very accurate like following image. (delta_adr = ADR(t) - ADR(t-1))\n![image](https://user-images.githubusercontent.com/24289602/128328750-1217fc5d-5280-42bd-be8b-0adf2485f3af.png)\n\nI used this to implement a forward hatch filter and a backward hatch filter and averaged them after applied Kaman smoothing.\n```python\n    prev_smooth = lat_lngs[0]\n    res = [prev_smooth]\n    for raw, delta_adr, delta_t in zip(\n        lat_lngs[1:, :], delta_adrs[1:, :], delta_ts[1:]\n    ):\n        if not all(np.isnan(delta_adr)):\n            smooth = 1 / M * raw + (M - 1) / M * (\n                prev_smooth + delta_adr * delta_t\n            )\n        else:\n            smooth = raw\n        prev_smooth = smooth\n        res.append(prev_smooth)\n   res = np.array(res)\n```\nThere is a parameter called M, so I optimized it using GroupKFold for each area. (Find the parameter that is the smallest in the train data and verify it with valid data.)\nSmoothing with this hatch filter was very effective.\nThere were many missing delta_adr values for some phones, so we filled them with values from other phones.\n\n# Weighted phone mean\n\nAfter applied hatch filter, I weighted average latitude and longitude using the weights of phone optimized by GroupKFold as well as M.\nSome of the `millisSinceGpsEpoch` had a little gap depending on the phone, so I used linear interpolation to fill it.\n\n\n# Snap2grid\n\nIn the downtown area, I performed a discrete optimization using linearly interpolated ground truth.\nTo prevent overfitting, the ground truth in itself was not included in the target to be snapped.\nThe objective function is as follows, and I used greedy search to find the optimal path.\n```python\ncost = distance2cur + distance2prev * alpha\n```\n\n# Stay point mean\n\nUsing the predicted value of the speed of lightgbm, I took the average for points below the threshold.\n\n# Ensemble\n\nUsing various parameters, I took the average of them and the CV was about 2.149.\nFinally, after adding other members' models and post-processing, we got CV  2.0569 and LB 2.285.\nIn Private, I knew that there was a lot of weight on highways and trees, so I decided on the ensemble weights taking them into account.",
    "1451405": "Good jod! i will learn from it!",
    "1454749": "Congrats  GrandMaster, and It was fun to work with you.\n\nOur almost solution is written by Shimacos and [monnu](https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/262195), so I will add some other point related shimacos's solution. (especially ADR)\n\n## ADR shift\n\nAccumulatedDeltaRange(ADR) is very strong about relative point relationship, but absolute point is a little weak.\nTherefore I calculate the difference between ADR and baseline's average location (after postprocess) in collection / phone, and shift to fit baseline's average.\nIt improved ADR's location performance.\n\n## Stop point correlation using ADR\n\nADR is also useful to stop point post processing.\nUsing not only speed(predict by lightgbm) but also ADR change from previous point make final submission improve.\n\n- When (speed < threshold1) & (adr_change < threshold2), stop flag is true\n- and there are more than N point in a row, these point replace to median point.\n\nThis is done after ensemble, and improve model."
  },
  "source": "meta"
}