{
  "id": 261774,
  "title": "18th Place Solution",
  "url": "/competitions/google-smartphone-decimeter-challenge/discussion/261774",
  "author_name": "kuto",
  "post_date": "2021-08-05T03:36:43.028000",
  "votes": 30,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Congratulations to all the winners, and thanks so much for hosting such an interesting competition! <br>\nThis competition is very tough for me, I'll share our solution.</p>\n<h2>Overview</h2>\n<p><img src=\"https://user-images.githubusercontent.com/43205304/128286772-0406d578-e9e0-4ea9-96df-3e0d30f2433c.png\" alt=\"GSDC-solution\"></p>\n<h2>1. Reproduce Baseline</h2>\n<p>In this competition, a baseline using WLS was shared(result file only). In order to reproduce this, we reproduced the baseline using derived files. It was difficult for us to reproduce the exact same results, but we were able to create baseline equivalent scores and decided to blend them with the original baseline file.</p>\n<p>We refered this notebook.<br>\n<a href=\"https://www.kaggle.com/hyperc/gsdc-reproducing-baseline-wls-on-one-measurement\" target=\"_blank\">https://www.kaggle.com/hyperc/gsdc-reproducing-baseline-wls-on-one-measurement</a></p>\n<p>Here are some of the things we did to create the baseline</p>\n<ul>\n<li>Use calculated satellite(GPS only) position by using OSR data(but effect is a little)</li>\n<li>Use only GPS/GALILEO/QZS data(other satellite is not good)</li>\n<li>Weight of WLS are hand-tuned (x**2 + 3x)</li>\n</ul>\n<p>After creating two new baseline locations using the above method, we blended them with the original baseline by taking a weighted average.</p>\n<h2>2. Area Classification</h2>\n<p>We automatically classified the collection into three categories as in the public version, and also defined a difficult area.<br>\nThe difficult area was defined as the area like following.</p>\n<ol>\n<li>Apply pre processing to train data</li>\n<li>Extract point which error is more than 5m</li>\n<li>Convert point to polygon by apply buffer to each point</li>\n<li>Define these polygon area as difficult area</li>\n</ol>\n<p>Defined difficult area are like this.<br>\n<img src=\"https://user-images.githubusercontent.com/43205304/128278785-0cb03ad5-45a9-4ed9-9203-2b2131cb7cef.png\" alt=\"スクリーンショット 2021-08-05 9 36 09\"></p>\n<p>This was used for snap to grid.</p>\n<h2>3. Pre/Post Processing</h2>\n<p>We used some shared notebook, Thanks you to the contributor.</p>\n<ul>\n<li><p>Outlier Correction<br>\n<a href=\"https://www.kaggle.com/dehokanta/baseline-post-processing-by-outlier-correction\" target=\"_blank\">https://www.kaggle.com/dehokanta/baseline-post-processing-by-outlier-correction</a></p></li>\n<li><p>Kalman Smoothing<br>\n<a href=\"https://www.kaggle.com/emaerthin/demonstration-of-the-kalman-filter\" target=\"_blank\">https://www.kaggle.com/emaerthin/demonstration-of-the-kalman-filter</a><br>\nWe apply linear interpolation to keep epoch width constant</p></li>\n<li><p>Phone Mean<br>\n<a href=\"https://www.kaggle.com/t88take/gsdc-phones-mean-prediction\" target=\"_blank\">https://www.kaggle.com/t88take/gsdc-phones-mean-prediction</a><br>\n<a href=\"https://www.kaggle.com/bpetrb/adaptive-gauss-phone-mean\" target=\"_blank\">https://www.kaggle.com/bpetrb/adaptive-gauss-phone-mean</a></p></li>\n<li><p>Remove Phone<br>\n<a href=\"https://www.kaggle.com/columbia2131/device-eda-interpolate-by-removing-device-en-ja\" target=\"_blank\">https://www.kaggle.com/columbia2131/device-eda-interpolate-by-removing-device-en-ja</a></p></li>\n<li><p>Snap to Grid<br>\nWe apply snap to grid only downtown area and difficult area.</p></li>\n<li><p>Stop Mean<br>\nI tried to take the average of the stopping points.<br>\nThe procedure is as follows.</p></li>\n</ul>\n<p>(1) Predict car speed by lightGBM</p>\n<ul>\n<li>target: speedMps in ground_truth.csv</li>\n<li>features<ul>\n<li>lag features with shift range is -30 ~ 30 (location, time, speed etc…)</li>\n<li>aggregate features</li></ul></li>\n</ul>\n<p>(2) If the prediction result is less than 0.95m/s and more than 2 consecutive seconds, make a group.<br>\n(3) Take the average for each group</p>\n<h2>4. Position estimation by imu data</h2>\n<p>Relative position correction was performed using IMU data for the only downtown area.<br>\nI refered this notebook. Thanks.<br>\n<a href=\"https://www.kaggle.com/alvinai9603/predict-next-point-with-the-imu-data\" target=\"_blank\">https://www.kaggle.com/alvinai9603/predict-next-point-with-the-imu-data</a></p>\n<ul>\n<li>model: lightGBM</li>\n<li>GroupKFold(group=\"phone\")</li>\n<li>features<ul>\n<li>lag features(shift range is -30~30)</li>\n<li>aggregate features(mean, std, max, min, median, skew, kart)</li></ul></li>\n</ul>",
  "messages": [
    {
      "id": 1450094,
      "postDate": "2021-08-05T03:36:43.030Z",
      "content": "<p>Congratulations to all the winners, and thanks so much for hosting such an interesting competition! <br>\nThis competition is very tough for me, I'll share our solution.</p>\n<h2>Overview</h2>\n<p><img src=\"https://user-images.githubusercontent.com/43205304/128286772-0406d578-e9e0-4ea9-96df-3e0d30f2433c.png\" alt=\"GSDC-solution\"></p>\n<h2>1. Reproduce Baseline</h2>\n<p>In this competition, a baseline using WLS was shared(result file only). In order to reproduce this, we reproduced the baseline using derived files. It was difficult for us to reproduce the exact same results, but we were able to create baseline equivalent scores and decided to blend them with the original baseline file.</p>\n<p>We refered this notebook.<br>\n<a href=\"https://www.kaggle.com/hyperc/gsdc-reproducing-baseline-wls-on-one-measurement\" target=\"_blank\">https://www.kaggle.com/hyperc/gsdc-reproducing-baseline-wls-on-one-measurement</a></p>\n<p>Here are some of the things we did to create the baseline</p>\n<ul>\n<li>Use calculated satellite(GPS only) position by using OSR data(but effect is a little)</li>\n<li>Use only GPS/GALILEO/QZS data(other satellite is not good)</li>\n<li>Weight of WLS are hand-tuned (x**2 + 3x)</li>\n</ul>\n<p>After creating two new baseline locations using the above method, we blended them with the original baseline by taking a weighted average.</p>\n<h2>2. Area Classification</h2>\n<p>We automatically classified the collection into three categories as in the public version, and also defined a difficult area.<br>\nThe difficult area was defined as the area like following.</p>\n<ol>\n<li>Apply pre processing to train data</li>\n<li>Extract point which error is more than 5m</li>\n<li>Convert point to polygon by apply buffer to each point</li>\n<li>Define these polygon area as difficult area</li>\n</ol>\n<p>Defined difficult area are like this.<br>\n<img src=\"https://user-images.githubusercontent.com/43205304/128278785-0cb03ad5-45a9-4ed9-9203-2b2131cb7cef.png\" alt=\"スクリーンショット 2021-08-05 9 36 09\"></p>\n<p>This was used for snap to grid.</p>\n<h2>3. Pre/Post Processing</h2>\n<p>We used some shared notebook, Thanks you to the contributor.</p>\n<ul>\n<li><p>Outlier Correction<br>\n<a href=\"https://www.kaggle.com/dehokanta/baseline-post-processing-by-outlier-correction\" target=\"_blank\">https://www.kaggle.com/dehokanta/baseline-post-processing-by-outlier-correction</a></p></li>\n<li><p>Kalman Smoothing<br>\n<a href=\"https://www.kaggle.com/emaerthin/demonstration-of-the-kalman-filter\" target=\"_blank\">https://www.kaggle.com/emaerthin/demonstration-of-the-kalman-filter</a><br>\nWe apply linear interpolation to keep epoch width constant</p></li>\n<li><p>Phone Mean<br>\n<a href=\"https://www.kaggle.com/t88take/gsdc-phones-mean-prediction\" target=\"_blank\">https://www.kaggle.com/t88take/gsdc-phones-mean-prediction</a><br>\n<a href=\"https://www.kaggle.com/bpetrb/adaptive-gauss-phone-mean\" target=\"_blank\">https://www.kaggle.com/bpetrb/adaptive-gauss-phone-mean</a></p></li>\n<li><p>Remove Phone<br>\n<a href=\"https://www.kaggle.com/columbia2131/device-eda-interpolate-by-removing-device-en-ja\" target=\"_blank\">https://www.kaggle.com/columbia2131/device-eda-interpolate-by-removing-device-en-ja</a></p></li>\n<li><p>Snap to Grid<br>\nWe apply snap to grid only downtown area and difficult area.</p></li>\n<li><p>Stop Mean<br>\nI tried to take the average of the stopping points.<br>\nThe procedure is as follows.</p></li>\n</ul>\n<p>(1) Predict car speed by lightGBM</p>\n<ul>\n<li>target: speedMps in ground_truth.csv</li>\n<li>features<ul>\n<li>lag features with shift range is -30 ~ 30 (location, time, speed etc…)</li>\n<li>aggregate features</li></ul></li>\n</ul>\n<p>(2) If the prediction result is less than 0.95m/s and more than 2 consecutive seconds, make a group.<br>\n(3) Take the average for each group</p>\n<h2>4. Position estimation by imu data</h2>\n<p>Relative position correction was performed using IMU data for the only downtown area.<br>\nI refered this notebook. Thanks.<br>\n<a href=\"https://www.kaggle.com/alvinai9603/predict-next-point-with-the-imu-data\" target=\"_blank\">https://www.kaggle.com/alvinai9603/predict-next-point-with-the-imu-data</a></p>\n<ul>\n<li>model: lightGBM</li>\n<li>GroupKFold(group=\"phone\")</li>\n<li>features<ul>\n<li>lag features(shift range is -30~30)</li>\n<li>aggregate features(mean, std, max, min, median, skew, kart)</li></ul></li>\n</ul>",
      "rawMarkdown": "Congratulations to all the winners, and thanks so much for hosting such an interesting competition! \nThis competition is very tough for me, I'll share our solution.\n\n## Overview\n![GSDC-solution](https://user-images.githubusercontent.com/43205304/128286772-0406d578-e9e0-4ea9-96df-3e0d30f2433c.png)\n\n\n\n## 1. Reproduce Baseline\nIn this competition, a baseline using WLS was shared(result file only). In order to reproduce this, we reproduced the baseline using derived files. It was difficult for us to reproduce the exact same results, but we were able to create baseline equivalent scores and decided to blend them with the original baseline file.\n\nWe refered this notebook.\nhttps://www.kaggle.com/hyperc/gsdc-reproducing-baseline-wls-on-one-measurement\n\nHere are some of the things we did to create the baseline\n- Use calculated satellite(GPS only) position by using OSR data(but effect is a little)\n- Use only GPS/GALILEO/QZS data(other satellite is not good)\n- Weight of WLS are hand-tuned (x**2 + 3x)\n\nAfter creating two new baseline locations using the above method, we blended them with the original baseline by taking a weighted average.\n\n## 2. Area Classification\nWe automatically classified the collection into three categories as in the public version, and also defined a difficult area.\nThe difficult area was defined as the area like following.\n1. Apply pre processing to train data\n2. Extract point which error is more than 5m\n3. Convert point to polygon by apply buffer to each point\n4. Define these polygon area as difficult area\n\nDefined difficult area are like this.\n![スクリーンショット 2021-08-05 9 36 09](https://user-images.githubusercontent.com/43205304/128278785-0cb03ad5-45a9-4ed9-9203-2b2131cb7cef.png)\n\nThis was used for snap to grid.\n\n## 3. Pre/Post Processing\nWe used some shared notebook, Thanks you to the contributor.\n\n- Outlier Correction\nhttps://www.kaggle.com/dehokanta/baseline-post-processing-by-outlier-correction\n\n- Kalman Smoothing\nhttps://www.kaggle.com/emaerthin/demonstration-of-the-kalman-filter\nWe apply linear interpolation to keep epoch width constant\n\n- Phone Mean\nhttps://www.kaggle.com/t88take/gsdc-phones-mean-prediction\nhttps://www.kaggle.com/bpetrb/adaptive-gauss-phone-mean\n\n- Remove Phone\nhttps://www.kaggle.com/columbia2131/device-eda-interpolate-by-removing-device-en-ja\n\n- Snap to Grid\nWe apply snap to grid only downtown area and difficult area.\n\n- Stop Mean\nI tried to take the average of the stopping points.\nThe procedure is as follows.\n\n(1) Predict car speed by lightGBM\n  - target: speedMps in ground_truth.csv\n  - features\n    - lag features with shift range is -30 ~ 30 (location, time, speed etc...)\n    - aggregate features\n\n(2) If the prediction result is less than 0.95m/s and more than 2 consecutive seconds, make a group.\n(3) Take the average for each group\n\n\n## 4. Position estimation by imu data\nRelative position correction was performed using IMU data for the only downtown area.\nI refered this notebook. Thanks.\nhttps://www.kaggle.com/alvinai9603/predict-next-point-with-the-imu-data\n\n- model: lightGBM\n- GroupKFold(group=\"phone\")\n- features\n  - lag features(shift range is -30~30)\n  - aggregate features(mean, std, max, min, median, skew, kart)\n\n",
      "votes": 30
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1450094": "Congratulations to all the winners, and thanks so much for hosting such an interesting competition! \nThis competition is very tough for me, I'll share our solution.\n\n## Overview\n![GSDC-solution](https://user-images.githubusercontent.com/43205304/128286772-0406d578-e9e0-4ea9-96df-3e0d30f2433c.png)\n\n\n\n## 1. Reproduce Baseline\nIn this competition, a baseline using WLS was shared(result file only). In order to reproduce this, we reproduced the baseline using derived files. It was difficult for us to reproduce the exact same results, but we were able to create baseline equivalent scores and decided to blend them with the original baseline file.\n\nWe refered this notebook.\nhttps://www.kaggle.com/hyperc/gsdc-reproducing-baseline-wls-on-one-measurement\n\nHere are some of the things we did to create the baseline\n- Use calculated satellite(GPS only) position by using OSR data(but effect is a little)\n- Use only GPS/GALILEO/QZS data(other satellite is not good)\n- Weight of WLS are hand-tuned (x**2 + 3x)\n\nAfter creating two new baseline locations using the above method, we blended them with the original baseline by taking a weighted average.\n\n## 2. Area Classification\nWe automatically classified the collection into three categories as in the public version, and also defined a difficult area.\nThe difficult area was defined as the area like following.\n1. Apply pre processing to train data\n2. Extract point which error is more than 5m\n3. Convert point to polygon by apply buffer to each point\n4. Define these polygon area as difficult area\n\nDefined difficult area are like this.\n![スクリーンショット 2021-08-05 9 36 09](https://user-images.githubusercontent.com/43205304/128278785-0cb03ad5-45a9-4ed9-9203-2b2131cb7cef.png)\n\nThis was used for snap to grid.\n\n## 3. Pre/Post Processing\nWe used some shared notebook, Thanks you to the contributor.\n\n- Outlier Correction\nhttps://www.kaggle.com/dehokanta/baseline-post-processing-by-outlier-correction\n\n- Kalman Smoothing\nhttps://www.kaggle.com/emaerthin/demonstration-of-the-kalman-filter\nWe apply linear interpolation to keep epoch width constant\n\n- Phone Mean\nhttps://www.kaggle.com/t88take/gsdc-phones-mean-prediction\nhttps://www.kaggle.com/bpetrb/adaptive-gauss-phone-mean\n\n- Remove Phone\nhttps://www.kaggle.com/columbia2131/device-eda-interpolate-by-removing-device-en-ja\n\n- Snap to Grid\nWe apply snap to grid only downtown area and difficult area.\n\n- Stop Mean\nI tried to take the average of the stopping points.\nThe procedure is as follows.\n\n(1) Predict car speed by lightGBM\n  - target: speedMps in ground_truth.csv\n  - features\n    - lag features with shift range is -30 ~ 30 (location, time, speed etc...)\n    - aggregate features\n\n(2) If the prediction result is less than 0.95m/s and more than 2 consecutive seconds, make a group.\n(3) Take the average for each group\n\n\n## 4. Position estimation by imu data\nRelative position correction was performed using IMU data for the only downtown area.\nI refered this notebook. Thanks.\nhttps://www.kaggle.com/alvinai9603/predict-next-point-with-the-imu-data\n\n- model: lightGBM\n- GroupKFold(group=\"phone\")\n- features\n  - lag features(shift range is -30~30)\n  - aggregate features(mean, std, max, min, median, skew, kart)\n\n"
  }
}