{
  "id": 239909,
  "title": "16th place solution",
  "url": "/competitions/indoor-location-navigation/writeups/exodia-reborn-at-motosumiyoshi-16th-place-solution",
  "author_name": "",
  "post_date": "2021-05-18T13:33:37.023Z",
  "votes": 35,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Congratulations to all the winners, and thanks so much for hosting such an interesting competition! I'll share our solution. </p>\n<h2>Summary</h2>\n<p><img src=\"https://user-images.githubusercontent.com/43205304/118599610-c62cff00-b7ea-11eb-9040-d49651249bbe.png\" alt=\"Indoor-solution (1)\"></p>\n<hr>\n<h2>dataset</h2>\n<p>Our dataset is based on <a href=\"https://www.kaggle.com/kokitanisaka/indoorunifiedwifids\" target=\"_blank\">indoor-unified-wifi-ds</a>. The major differences are as follows.</p>\n<ul>\n<li>Remove the data, if the difference between timestamp of waypoint and last seen timestamp of wifi is more than 10s. <strong>(This is very important)</strong></li>\n<li>The test data was also created as a wifi-based dataset in the same way as the train in order to average the timestamp during inference. So there are about three times as many lines as the original.</li>\n<li>The hidden waypoints are complemented by <a href=\"https://www.kaggle.com/kuto0633/linear-interpolation-for-waypoint-in-wifi-dataset\" target=\"_blank\">Linear Interpolation</a> and <a href=\"https://www.kaggle.com/arnaudcapitaine/get-indoor-location-by-means-of-kalman-smoother\" target=\"_blank\">Kalman Filter</a>. This can be expected to have a padding effect on the data.<br>\n<img alt=\"スクリーンショット 2021-05-18 8 41 25\" src=\"https://user-images.githubusercontent.com/43205304/118569421-e2ae4480-b7b4-11eb-9b6f-713566812365.png\"></li>\n</ul>\n<hr>\n<h2>model</h2>\n<h3>model for xy</h3>\n<p>Our Model is based on <a href=\"https://www.kaggle.com/kokitanisaka/lstm-by-keras-with-unified-wi-fi-feats\" target=\"_blank\">LSTM by Keras with Unified Wi-Fi Feats</a>. Since this LSTM model is not time series based, we also use the MLP model. <br>\nUsing features is following.</p>\n<ul>\n<li>80 pieces BSSID</li>\n<li>80 pieces RSSI</li>\n<li>site id </li>\n<li>floor (In inference, we use values which other model predict)</li>\n</ul>\n<h3>model for floor</h3>\n<ul>\n<li>lightGBM (also try BiLSTM)</li>\n<li>Learning as a classification</li>\n<li>Create model in each site</li>\n<li>Create dataset each path file as one line</li>\n<li>StratifiedGroupKFold</li>\n<li>features<ul>\n<li>BSSID  </li>\n<li>Mean RSSI</li>\n<li>Max RSSI</li>\n<li>Extract BSSIDs that are present in both test and train.</li></ul></li>\n</ul>\n<hr>\n<h2>training (2stage)</h2>\n<h3>1st stage</h3>\n<ul>\n<li>Almost as shown in the figure above.</li>\n<li>Custom loss<br>\nEach data in our dataset has timediff(time difference between timestamp of waypoint and timestamp of wifi group).  The larger this timediff is, the larger the discrepancy between the given target and the ground truth. So we apply Weighted-Loss according to timediff.</li>\n</ul>\n<pre><code># timediff -&gt; weight\ntimediff = df['timediff'].astype(np.float32).abs().values\nweight = 1- (timediff/np.max(timediff)) \n\nclass WeightedMSELoss(nn.Module):\n    def __init__(self):\n        super().__init__()\n        self.loss = nn.MSELoss(reduction='none')\n\n    def forward(self, input, target, weight):\n        input = input.float()\n        target = target.float()\n        weight = torch.stack((weight, weight), 1).float()  \n        loss = self.loss(input, target) * weight \n        return loss.mean()\n</code></pre>\n<h3>2nd stage</h3>\n<p>Re-learn by adding the following elements.</p>\n<ul>\n<li>Add test data by pseudo labelig </li>\n<li>Remove train data  if oof's error is over 40m (also try to remove over 20m)</li>\n</ul>\n<hr>\n<h2>post processing (pp)</h2>\n<p>We use three-pp which were shared in Notebook &amp; Discussion.<br>\nWe repeated three-pp <strong>6 times</strong>. It is very effective.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization\" target=\"_blank\">Cost minimaization</a><br>\nThere are difference between delta by sensor and delta by target. So we corrected sensor delta by the coefficients which obtained by linear regression  in each site and each floor.</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th></th>\n<th></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><img alt=\"スクリーンショット 2021-05-18 8 49 51\" src=\"https://user-images.githubusercontent.com/43205304/118569982-16d63500-b7b6-11eb-8f70-35f1c2c10c1c.png\"></td>\n<td><img alt=\"スクリーンショット 2021-05-18 8 49 59\" src=\"https://user-images.githubusercontent.com/43205304/118569988-189ff880-b7b6-11eb-9e0e-060ece210127.png\"></td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li><a href=\"https://www.kaggle.com/robikscube/indoor-navigation-snap-to-grid-post-processing\" target=\"_blank\">Snap to grid</a><br>\nWe automatically generated 3 pattern extra grids instead of hand-labeling grids like following. </li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>sparse</th>\n<th>dense</th>\n<th>edge</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><img alt=\"スクリーンショット 2021-05-18 8 13 53\" src=\"https://user-images.githubusercontent.com/43205304/118567646-1dae7900-b7b1-11eb-9425-ffb414c7f923.png\"></td>\n<td><img alt=\"スクリーンショット 2021-05-18 8 14 27\" src=\"https://user-images.githubusercontent.com/43205304/118567717-42a2ec00-b7b1-11eb-957f-b18649ea3f0d.png\"></td>\n<td><img alt=\"スクリーンショット 2021-05-18 8 14 07\" src=\"https://user-images.githubusercontent.com/43205304/118567652-21420000-b7b1-11eb-92c3-29329d2d1fd5.png\"></td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li><a href=\"https://www.kaggle.com/iwatatakuya/use-leakage-considering-device-id-postprocess\" target=\"_blank\">device id leak</a></li>\n</ul>\n<h2>ensemble</h2>\n<p>We did stacking and extra grid ensemble. We did post processing with 4 different patterns. Then, did ensemble by weighted average.</p>\n<p>4 pattern is here.<br>\n① snap to grid’s threshold=None / sparse extra grid<br>\n② snap to grid’s threshold=None / dense  extra grid<br>\n③ snap to grid’s threshold=None / edge  extra grid<br>\n④ snap to grid’s threshold=5 / only train grid</p>\n<hr>\n<p>Thank you!</p>",
  "messages": [
    {
      "id": "1312381",
      "postDate": "05/18/2021 02:53:18",
      "content": "<p>Congratulations to all the winners, and thanks so much for hosting such an interesting competition! I'll share our solution. </p>\n<h2>Summary</h2>\n<p><img src=\"https://user-images.githubusercontent.com/43205304/118599610-c62cff00-b7ea-11eb-9040-d49651249bbe.png\" alt=\"Indoor-solution (1)\"></p>\n<hr>\n<h2>dataset</h2>\n<p>Our dataset is based on <a href=\"https://www.kaggle.com/kokitanisaka/indoorunifiedwifids\" target=\"_blank\">indoor-unified-wifi-ds</a>. The major differences are as follows.</p>\n<ul>\n<li>Remove the data, if the difference between timestamp of waypoint and last seen timestamp of wifi is more than 10s. <strong>(This is very important)</strong></li>\n<li>The test data was also created as a wifi-based dataset in the same way as the train in order to average the timestamp during inference. So there are about three times as many lines as the original.</li>\n<li>The hidden waypoints are complemented by <a href=\"https://www.kaggle.com/kuto0633/linear-interpolation-for-waypoint-in-wifi-dataset\" target=\"_blank\">Linear Interpolation</a> and <a href=\"https://www.kaggle.com/arnaudcapitaine/get-indoor-location-by-means-of-kalman-smoother\" target=\"_blank\">Kalman Filter</a>. This can be expected to have a padding effect on the data.<br>\n<img alt=\"スクリーンショット 2021-05-18 8 41 25\" src=\"https://user-images.githubusercontent.com/43205304/118569421-e2ae4480-b7b4-11eb-9b6f-713566812365.png\"></li>\n</ul>\n<hr>\n<h2>model</h2>\n<h3>model for xy</h3>\n<p>Our Model is based on <a href=\"https://www.kaggle.com/kokitanisaka/lstm-by-keras-with-unified-wi-fi-feats\" target=\"_blank\">LSTM by Keras with Unified Wi-Fi Feats</a>. Since this LSTM model is not time series based, we also use the MLP model. <br>\nUsing features is following.</p>\n<ul>\n<li>80 pieces BSSID</li>\n<li>80 pieces RSSI</li>\n<li>site id </li>\n<li>floor (In inference, we use values which other model predict)</li>\n</ul>\n<h3>model for floor</h3>\n<ul>\n<li>lightGBM (also try BiLSTM)</li>\n<li>Learning as a classification</li>\n<li>Create model in each site</li>\n<li>Create dataset each path file as one line</li>\n<li>StratifiedGroupKFold</li>\n<li>features<ul>\n<li>BSSID  </li>\n<li>Mean RSSI</li>\n<li>Max RSSI</li>\n<li>Extract BSSIDs that are present in both test and train.</li></ul></li>\n</ul>\n<hr>\n<h2>training (2stage)</h2>\n<h3>1st stage</h3>\n<ul>\n<li>Almost as shown in the figure above.</li>\n<li>Custom loss<br>\nEach data in our dataset has timediff(time difference between timestamp of waypoint and timestamp of wifi group).  The larger this timediff is, the larger the discrepancy between the given target and the ground truth. So we apply Weighted-Loss according to timediff.</li>\n</ul>\n<pre><code># timediff -&gt; weight\ntimediff = df['timediff'].astype(np.float32).abs().values\nweight = 1- (timediff/np.max(timediff)) \n\nclass WeightedMSELoss(nn.Module):\n    def __init__(self):\n        super().__init__()\n        self.loss = nn.MSELoss(reduction='none')\n\n    def forward(self, input, target, weight):\n        input = input.float()\n        target = target.float()\n        weight = torch.stack((weight, weight), 1).float()  \n        loss = self.loss(input, target) * weight \n        return loss.mean()\n</code></pre>\n<h3>2nd stage</h3>\n<p>Re-learn by adding the following elements.</p>\n<ul>\n<li>Add test data by pseudo labelig </li>\n<li>Remove train data  if oof's error is over 40m (also try to remove over 20m)</li>\n</ul>\n<hr>\n<h2>post processing (pp)</h2>\n<p>We use three-pp which were shared in Notebook &amp; Discussion.<br>\nWe repeated three-pp <strong>6 times</strong>. It is very effective.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization\" target=\"_blank\">Cost minimaization</a><br>\nThere are difference between delta by sensor and delta by target. So we corrected sensor delta by the coefficients which obtained by linear regression  in each site and each floor.</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th></th>\n<th></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><img alt=\"スクリーンショット 2021-05-18 8 49 51\" src=\"https://user-images.githubusercontent.com/43205304/118569982-16d63500-b7b6-11eb-8f70-35f1c2c10c1c.png\"></td>\n<td><img alt=\"スクリーンショット 2021-05-18 8 49 59\" src=\"https://user-images.githubusercontent.com/43205304/118569988-189ff880-b7b6-11eb-9e0e-060ece210127.png\"></td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li><a href=\"https://www.kaggle.com/robikscube/indoor-navigation-snap-to-grid-post-processing\" target=\"_blank\">Snap to grid</a><br>\nWe automatically generated 3 pattern extra grids instead of hand-labeling grids like following. </li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>sparse</th>\n<th>dense</th>\n<th>edge</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><img alt=\"スクリーンショット 2021-05-18 8 13 53\" src=\"https://user-images.githubusercontent.com/43205304/118567646-1dae7900-b7b1-11eb-9425-ffb414c7f923.png\"></td>\n<td><img alt=\"スクリーンショット 2021-05-18 8 14 27\" src=\"https://user-images.githubusercontent.com/43205304/118567717-42a2ec00-b7b1-11eb-957f-b18649ea3f0d.png\"></td>\n<td><img alt=\"スクリーンショット 2021-05-18 8 14 07\" src=\"https://user-images.githubusercontent.com/43205304/118567652-21420000-b7b1-11eb-92c3-29329d2d1fd5.png\"></td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li><a href=\"https://www.kaggle.com/iwatatakuya/use-leakage-considering-device-id-postprocess\" target=\"_blank\">device id leak</a></li>\n</ul>\n<h2>ensemble</h2>\n<p>We did stacking and extra grid ensemble. We did post processing with 4 different patterns. Then, did ensemble by weighted average.</p>\n<p>4 pattern is here.<br>\n① snap to grid’s threshold=None / sparse extra grid<br>\n② snap to grid’s threshold=None / dense  extra grid<br>\n③ snap to grid’s threshold=None / edge  extra grid<br>\n④ snap to grid’s threshold=5 / only train grid</p>\n<hr>\n<p>Thank you!</p>",
      "rawMarkdown": "Congratulations to all the winners, and thanks so much for hosting such an interesting competition! I'll share our solution. \n\n## Summary\n![Indoor-solution (1)](https://user-images.githubusercontent.com/43205304/118599610-c62cff00-b7ea-11eb-9040-d49651249bbe.png)\n\n\n---\n\n## dataset\nOur dataset is based on [indoor-unified-wifi-ds](https://www.kaggle.com/kokitanisaka/indoorunifiedwifids). The major differences are as follows.\n- Remove the data, if the difference between timestamp of waypoint and last seen timestamp of wifi is more than 10s. **(This is very important)**\n- The test data was also created as a wifi-based dataset in the same way as the train in order to average the timestamp during inference. So there are about three times as many lines as the original.\n- The hidden waypoints are complemented by [Linear Interpolation](https://www.kaggle.com/kuto0633/linear-interpolation-for-waypoint-in-wifi-dataset) and [Kalman Filter](https://www.kaggle.com/arnaudcapitaine/get-indoor-location-by-means-of-kalman-smoother). This can be expected to have a padding effect on the data.\n<img width=\"320\" alt=\"スクリーンショット 2021-05-18 8 41 25\" src=\"https://user-images.githubusercontent.com/43205304/118569421-e2ae4480-b7b4-11eb-9b6f-713566812365.png\">\n\n---\n## model\n### model for xy\nOur Model is based on [LSTM by Keras with Unified Wi-Fi Feats](https://www.kaggle.com/kokitanisaka/lstm-by-keras-with-unified-wi-fi-feats). Since this LSTM model is not time series based, we also use the MLP model. \nUsing features is following.\n- 80 pieces BSSID\n- 80 pieces RSSI\n- site id \n- floor (In inference, we use values which other model predict)\n\n###  model for floor\n- lightGBM (also try BiLSTM)\n- Learning as a classification\n- Create model in each site\n- Create dataset each path file as one line\n- StratifiedGroupKFold\n- features\n    - BSSID  \n    - Mean RSSI\n    - Max RSSI\n    - Extract BSSIDs that are present in both test and train.\n\n\n---\n## training (2stage) \n\n### 1st stage\n- Almost as shown in the figure above.\n- Custom loss\nEach data in our dataset has timediff(time difference between timestamp of waypoint and timestamp of wifi group).  The larger this timediff is, the larger the discrepancy between the given target and the ground truth. So we apply Weighted-Loss according to timediff.\n\n```python\n# timediff -> weight\ntimediff = df['timediff'].astype(np.float32).abs().values\nweight = 1- (timediff/np.max(timediff)) \n\nclass WeightedMSELoss(nn.Module):\n    def __init__(self):\n        super().__init__()\n        self.loss = nn.MSELoss(reduction='none')\n\n    def forward(self, input, target, weight):\n        input = input.float()\n        target = target.float()\n        weight = torch.stack((weight, weight), 1).float()  \n        loss = self.loss(input, target) * weight \n        return loss.mean()\n```\n\n### 2nd stage\nRe-learn by adding the following elements.\n- Add test data by pseudo labelig \n- Remove train data  if oof's error is over 40m (also try to remove over 20m)\n\n---\n## post processing (pp)\n\nWe use three-pp which were shared in Notebook & Discussion.\nWe repeated three-pp **6 times**. It is very effective.\n\n- [Cost minimaization](https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization)\nThere are difference between delta by sensor and delta by target. So we corrected sensor delta by the coefficients which obtained by linear regression  in each site and each floor.\n\n| | |\n| --- | --- |\n| <img width=\"362\" alt=\"スクリーンショット 2021-05-18 8 49 51\" src=\"https://user-images.githubusercontent.com/43205304/118569982-16d63500-b7b6-11eb-8f70-35f1c2c10c1c.png\"> | <img width=\"361\" alt=\"スクリーンショット 2021-05-18 8 49 59\" src=\"https://user-images.githubusercontent.com/43205304/118569988-189ff880-b7b6-11eb-9e0e-060ece210127.png\">  |\n  \n  \n- [Snap to grid](https://www.kaggle.com/robikscube/indoor-navigation-snap-to-grid-post-processing)\nWe automatically generated 3 pattern extra grids instead of hand-labeling grids like following. \n  \n| sparse | dense | edge |\n| --- | --- | --- |\n| <img width=\"495\" alt=\"スクリーンショット 2021-05-18 8 13 53\" src=\"https://user-images.githubusercontent.com/43205304/118567646-1dae7900-b7b1-11eb-9425-ffb414c7f923.png\"> | <img width=\"495\" alt=\"スクリーンショット 2021-05-18 8 14 27\" src=\"https://user-images.githubusercontent.com/43205304/118567717-42a2ec00-b7b1-11eb-957f-b18649ea3f0d.png\"> | <img width=\"492\" alt=\"スクリーンショット 2021-05-18 8 14 07\" src=\"https://user-images.githubusercontent.com/43205304/118567652-21420000-b7b1-11eb-92c3-29329d2d1fd5.png\">\n   \n   \n- [device id leak](https://www.kaggle.com/iwatatakuya/use-leakage-considering-device-id-postprocess)\n\n## ensemble\nWe did stacking and extra grid ensemble. We did post processing with 4 different patterns. Then, did ensemble by weighted average.\n\n4 pattern is here.\n① snap to grid’s threshold=None / sparse extra grid\n② snap to grid’s threshold=None / dense  extra grid\n③ snap to grid’s threshold=None / edge  extra grid\n④ snap to grid’s threshold=5 / only train grid\n\n---\n\nThank you!",
      "votes": null
    },
    {
      "id": "1312429",
      "postDate": "05/18/2021 03:52:10",
      "content": "<p>Thank you for publishing a great solution.Can you tell me about the post-processing repeat?</p>\n<p>I have tried repeats and found them to be effective in some cases and vice versa. I wasn't confident about the effectiveness of repeat so I didn't use it in my final sub.</p>\n<p>What was your score before post-processing?<br>\n Also, how effective were the repeats without extra grids?</p>\n<p>Personally, I think that if there are no extra grid points, there will be no effect of repeating the process.</p>",
      "rawMarkdown": "Thank you for publishing a great solution.Can you tell me about the post-processing repeat?\n\nI have tried repeats and found them to be effective in some cases and vice versa. I wasn't confident about the effectiveness of repeat so I didn't use it in my final sub.\n\nWhat was your score before post-processing?\n Also, how effective were the repeats without extra grids?\n\nPersonally, I think that if there are no extra grid points, there will be no effect of repeating the process.",
      "votes": null
    },
    {
      "id": "1312465",
      "postDate": "05/18/2021 04:41:31",
      "content": "<p><a href=\"https://www.kaggle.com/dehokanta\" target=\"_blank\">@dehokanta</a> Thanks for your comment.</p>\n<blockquote>\n  <p>What was your score before post-processing?<br>\n  Also, how effective were the repeats without extra grids?</p>\n</blockquote>\n<p>Our single model's CV is here. <br>\nThis score is used extra grid, but the one which is not used extra grid is almost same improvement.<br>\nAfter the 3 times, the effect of pp slowed down but the CV increased, so we decided 6 times.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>CV score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CV(not pp)</td>\n<td>6.311</td>\n</tr>\n<tr>\n<td>CV(pp×1)</td>\n<td>4.672</td>\n</tr>\n<tr>\n<td>CV(pp×2)</td>\n<td>4.377</td>\n</tr>\n</tbody>\n</table>\n<blockquote>\n  <p>Personally, I think that if there are no extra grid points, there will be no effect of repeating the process.</p>\n</blockquote>\n<p>There are also good effect by repeating without extra-grids in our case. By doing snap to grid(N-1 times.), the predicted waypoints closer to ground truth, so the result of Cost Minimaization(N times) also changed, and the score improved, I think.</p>",
      "rawMarkdown": "dehokanta Thanks for your comment.\n\n> What was your score before post-processing?\nAlso, how effective were the repeats without extra grids?\n\nOur single model's CV is here. \nThis score is used extra grid, but the one which is not used extra grid is almost same improvement.\nAfter the 3 times, the effect of pp slowed down but the CV increased, so we decided 6 times.\n\n| | CV score | \n| --- | --- | \n| CV(not pp) | 6.311 |\n| CV(pp×1) | 4.672 |  \n| CV(pp×2) | 4.377 |  \n\n> Personally, I think that if there are no extra grid points, there will be no effect of repeating the process.\n\nThere are also good effect by repeating without extra-grids in our case. By doing snap to grid(N-1 times.), the predicted waypoints closer to ground truth, so the result of Cost Minimaization(N times) also changed, and the score improved, I think.",
      "votes": null
    },
    {
      "id": "1312549",
      "postDate": "05/18/2021 05:37:52",
      "content": "<p><a href=\"https://www.kaggle.com/kuto0633\" target=\"_blank\">@kuto0633</a> Thanks for the quick reply.<br>\nThe result of our PP (without hand label) loop was this.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Public Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>without PP</td>\n<td>5.096</td>\n</tr>\n<tr>\n<td>PP×1</td>\n<td>3.569</td>\n</tr>\n<tr>\n<td>PP×4</td>\n<td>3.766</td>\n</tr>\n</tbody>\n</table>\n<p>I didn't use any additional grids, so Snap to Grid moved  from the correct position to the wrong position, which may be why our score didn't improve.</p>\n<p>Thank you very much for valuable comments.</p>",
      "rawMarkdown": "kuto0633 Thanks for the quick reply.\nThe result of our PP (without hand label) loop was this.\n|  | Public Score |\n| --- | --- |\n|  without PP | 5.096 |\n|  PP×1 | 3.569 |\n|  PP×4 | 3.766 |\n\nI didn't use any additional grids, so Snap to Grid moved  from the correct position to the wrong position, which may be why our score didn't improve.\n\nThank you very much for valuable comments.",
      "votes": null
    },
    {
      "id": "1313046",
      "postDate": "05/18/2021 11:37:04",
      "content": "<p>Thank you for sharing this great solution!<br>\nI learned a lot from this 😃 (eg. Custom loss, Removing big timediff samples or large loss samples, pseudo labeling)</p>\n<p>I have one question about your model.</p>\n<blockquote>\n  <p>Since this LSTM model is not time series based, we also use the MLP model.</p>\n</blockquote>\n<p>Does this mean you created two different models (not time-series lstm and mlp) and the latter MLP model used time-series information like gru or lstm ?</p>\n<p>Thanks in advance!</p>",
      "rawMarkdown": "Thank you for sharing this great solution!\nI learned a lot from this 😃 (eg. Custom loss, Removing big timediff samples or large loss samples, pseudo labeling)\n\nI have one question about your model.\n> Since this LSTM model is not time series based, we also use the MLP model.\n\nDoes this mean you created two different models (not time-series lstm and mlp) and the latter MLP model used time-series information like gru or lstm ?\n\nThanks in advance!",
      "votes": null
    },
    {
      "id": "1313076",
      "postDate": "05/18/2021 11:52:40",
      "content": "<p><a href=\"https://www.kaggle.com/yutoshibata\" target=\"_blank\">@yutoshibata</a> Thanks! I'm glad such your comments.</p>\n<blockquote>\n  <p>you created two different models (not time-series lstm and mlp) </p>\n</blockquote>\n<p>Yes, I and another teammate created LSTM, and the other created MLP.</p>\n<blockquote>\n  <p>the latter MLP model used time-series information like gru or lstm ?</p>\n</blockquote>\n<p>No,  I mean that our LSTM is almost the same as MLP because it doesn't deal with time series. So our any model are not time-series.</p>",
      "rawMarkdown": "yutoshibata Thanks! I'm glad such your comments.\n> you created two different models (not time-series lstm and mlp) \n\nYes, I and another teammate created LSTM, and the other created MLP.\n\n> the latter MLP model used time-series information like gru or lstm ?\n\nNo,  I mean that our LSTM is almost the same as MLP because it doesn't deal with time series. So our any model are not time-series.",
      "votes": null
    },
    {
      "id": "1313315",
      "postDate": "05/18/2021 14:03:36",
      "content": "<p><a href=\"https://www.kaggle.com/kuto0633\" target=\"_blank\">@kuto0633</a> Thanks ! I totally understand your comment 😃<br>\nI used MLP too but CV score was 7.3….<br>\nI think removing some samples or using KF or custom loss made this big difference.<br>\nAnyway, thank you for explaining your solution kindly🙌</p>",
      "rawMarkdown": "kuto0633 Thanks ! I totally understand your comment 😃\nI used MLP too but CV score was 7.3....\nI think removing some samples or using KF or custom loss made this big difference.\nAnyway, thank you for explaining your solution kindly🙌",
      "votes": null
    },
    {
      "id": "1313406",
      "postDate": "05/18/2021 14:43:20",
      "content": "<p>Thank you for sharing!<br>\nI had a lot to learn from your great solutions.</p>\n<blockquote>\n  <p>Remove the data, if the difference between timestamp of waypoint and last seen timestamp of wifi is more than 10s. (This is very important)  </p>\n</blockquote>\n<p>Actually, I was able to improve my score significantly by adopting this idea as well.<br>\nHowever, my approach to discarding data is a little different from yours. I did not give any superiority to the importance based on the time difference from the waypoint, so I focused on the relation between last seen timestamp and wifi timestamp at the same row in the text file. For example, </p>\n<p>timestamp=1000, bssid=AAA, lastseen_ts=950 &lt;-- keep<br>\ntimestamp=1200, bssid=AAA, lastseen_ts=950 &lt;-- discard because this information is older than the previous row</p>\n<p>Now that I know how to skillfully utilize the time difference from the waypoint, I feel that your method is better.<br>\nI would be happy to hear your frank opinion.</p>",
      "rawMarkdown": "Thank you for sharing!\nI had a lot to learn from your great solutions.\n> Remove the data, if the difference between timestamp of waypoint and last seen timestamp of wifi is more than 10s. (This is very important)  \n\nActually, I was able to improve my score significantly by adopting this idea as well.\nHowever, my approach to discarding data is a little different from yours. I did not give any superiority to the importance based on the time difference from the waypoint, so I focused on the relation between last seen timestamp and wifi timestamp at the same row in the text file. For example, \n  \ntimestamp=1000, bssid=AAA, lastseen_ts=950 <-- keep\ntimestamp=1200, bssid=AAA, lastseen_ts=950 <-- discard because this information is older than the previous row\n  \nNow that I know how to skillfully utilize the time difference from the waypoint, I feel that your method is better.\nI would be happy to hear your frank opinion.",
      "votes": null
    },
    {
      "id": "1313548",
      "postDate": "05/18/2021 15:47:10",
      "content": "<p><a href=\"https://www.kaggle.com/horsek\" target=\"_blank\">@horsek</a> Thanks and congrats your silver medal!<br>\nI understand your approach. Maybe it almost same effect as ours if you set the appropriate timediff-threshold(=10s in our case).</p>\n<pre><code>t_l: wifi last seen timestamp\nt_g: wifi group timestamp\nt_w: waypoint timestamp\n\nour timediff\n = t_w - t_l\n\nyour timediff\n= t_g - t_l \n= (t_g - t_l) + (t_w - t_w) \n= (t_w - t_l) + (t_g - t_w) \n= (our timediff) + (t_g - t_w)\n</code></pre>\n<p>where (t_g - t_w) should be constant if we and you use kouki's dataset.</p>\n<p>If there are any mistakes, please point them out.</p>",
      "rawMarkdown": "horsek Thanks and congrats your silver medal!\nI understand your approach. Maybe it almost same effect as ours if you set the appropriate timediff-threshold(=10s in our case).\n\n```\nt_l: wifi last seen timestamp\nt_g: wifi group timestamp\nt_w: waypoint timestamp\n\nour timediff\n = t_w - t_l\n\nyour timediff\n= t_g - t_l \n= (t_g - t_l) + (t_w - t_w) \n= (t_w - t_l) + (t_g - t_w) \n= (our timediff) + (t_g - t_w)\n\n```\nwhere (t_g - t_w) should be constant if we and you use kouki's dataset.\n\nIf there are any mistakes, please point them out.",
      "votes": null
    },
    {
      "id": "1313720",
      "postDate": "05/18/2021 17:14:59",
      "content": "<p>Thank you! Congrats to you too on your silver medal.<br>\nI see. I agree that we are essentially doing the same thing. And the fact that both of our models have shown pretty good improvement tells that you are right.</p>",
      "rawMarkdown": "Thank you! Congrats to you too on your silver medal.\nI see. I agree that we are essentially doing the same thing. And the fact that both of our models have shown pretty good improvement tells that you are right.",
      "votes": null
    },
    {
      "id": "1323375",
      "postDate": "05/26/2021 07:23:23",
      "content": "<p>Thanks for Sharing this to us. <br>\nUpvote for you. </p>",
      "rawMarkdown": "Thanks for Sharing this to us. \nUpvote for you.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1312429,
      "author_name": "dehokanta",
      "author_url": "",
      "post_date": "05/18/2021 03:52:10",
      "content": "<p>Thank you for publishing a great solution.Can you tell me about the post-processing repeat?</p>\n<p>I have tried repeats and found them to be effective in some cases and vice versa. I wasn't confident about the effectiveness of repeat so I didn't use it in my final sub.</p>\n<p>What was your score before post-processing?<br>\n Also, how effective were the repeats without extra grids?</p>\n<p>Personally, I think that if there are no extra grid points, there will be no effect of repeating the process.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1312465,
          "author_name": "kuto0633",
          "author_url": "",
          "post_date": "05/18/2021 04:41:31",
          "content": "<p><a href=\"https://www.kaggle.com/dehokanta\" target=\"_blank\">@dehokanta</a> Thanks for your comment.</p>\n<blockquote>\n  <p>What was your score before post-processing?<br>\n  Also, how effective were the repeats without extra grids?</p>\n</blockquote>\n<p>Our single model's CV is here. <br>\nThis score is used extra grid, but the one which is not used extra grid is almost same improvement.<br>\nAfter the 3 times, the effect of pp slowed down but the CV increased, so we decided 6 times.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>CV score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CV(not pp)</td>\n<td>6.311</td>\n</tr>\n<tr>\n<td>CV(pp×1)</td>\n<td>4.672</td>\n</tr>\n<tr>\n<td>CV(pp×2)</td>\n<td>4.377</td>\n</tr>\n</tbody>\n</table>\n<blockquote>\n  <p>Personally, I think that if there are no extra grid points, there will be no effect of repeating the process.</p>\n</blockquote>\n<p>There are also good effect by repeating without extra-grids in our case. By doing snap to grid(N-1 times.), the predicted waypoints closer to ground truth, so the result of Cost Minimaization(N times) also changed, and the score improved, I think.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1312549,
          "author_name": "dehokanta",
          "author_url": "",
          "post_date": "05/18/2021 05:37:52",
          "content": "<p><a href=\"https://www.kaggle.com/kuto0633\" target=\"_blank\">@kuto0633</a> Thanks for the quick reply.<br>\nThe result of our PP (without hand label) loop was this.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Public Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>without PP</td>\n<td>5.096</td>\n</tr>\n<tr>\n<td>PP×1</td>\n<td>3.569</td>\n</tr>\n<tr>\n<td>PP×4</td>\n<td>3.766</td>\n</tr>\n</tbody>\n</table>\n<p>I didn't use any additional grids, so Snap to Grid moved  from the correct position to the wrong position, which may be why our score didn't improve.</p>\n<p>Thank you very much for valuable comments.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1313046,
      "author_name": "yutoshibata",
      "author_url": "",
      "post_date": "05/18/2021 11:37:04",
      "content": "<p>Thank you for sharing this great solution!<br>\nI learned a lot from this 😃 (eg. Custom loss, Removing big timediff samples or large loss samples, pseudo labeling)</p>\n<p>I have one question about your model.</p>\n<blockquote>\n  <p>Since this LSTM model is not time series based, we also use the MLP model.</p>\n</blockquote>\n<p>Does this mean you created two different models (not time-series lstm and mlp) and the latter MLP model used time-series information like gru or lstm ?</p>\n<p>Thanks in advance!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1313076,
          "author_name": "kuto0633",
          "author_url": "",
          "post_date": "05/18/2021 11:52:40",
          "content": "<p><a href=\"https://www.kaggle.com/yutoshibata\" target=\"_blank\">@yutoshibata</a> Thanks! I'm glad such your comments.</p>\n<blockquote>\n  <p>you created two different models (not time-series lstm and mlp) </p>\n</blockquote>\n<p>Yes, I and another teammate created LSTM, and the other created MLP.</p>\n<blockquote>\n  <p>the latter MLP model used time-series information like gru or lstm ?</p>\n</blockquote>\n<p>No,  I mean that our LSTM is almost the same as MLP because it doesn't deal with time series. So our any model are not time-series.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1313315,
          "author_name": "yutoshibata",
          "author_url": "",
          "post_date": "05/18/2021 14:03:36",
          "content": "<p><a href=\"https://www.kaggle.com/kuto0633\" target=\"_blank\">@kuto0633</a> Thanks ! I totally understand your comment 😃<br>\nI used MLP too but CV score was 7.3….<br>\nI think removing some samples or using KF or custom loss made this big difference.<br>\nAnyway, thank you for explaining your solution kindly🙌</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1313406,
      "author_name": "horsek",
      "author_url": "",
      "post_date": "05/18/2021 14:43:20",
      "content": "<p>Thank you for sharing!<br>\nI had a lot to learn from your great solutions.</p>\n<blockquote>\n  <p>Remove the data, if the difference between timestamp of waypoint and last seen timestamp of wifi is more than 10s. (This is very important)  </p>\n</blockquote>\n<p>Actually, I was able to improve my score significantly by adopting this idea as well.<br>\nHowever, my approach to discarding data is a little different from yours. I did not give any superiority to the importance based on the time difference from the waypoint, so I focused on the relation between last seen timestamp and wifi timestamp at the same row in the text file. For example, </p>\n<p>timestamp=1000, bssid=AAA, lastseen_ts=950 &lt;-- keep<br>\ntimestamp=1200, bssid=AAA, lastseen_ts=950 &lt;-- discard because this information is older than the previous row</p>\n<p>Now that I know how to skillfully utilize the time difference from the waypoint, I feel that your method is better.<br>\nI would be happy to hear your frank opinion.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1313548,
          "author_name": "kuto0633",
          "author_url": "",
          "post_date": "05/18/2021 15:47:10",
          "content": "<p><a href=\"https://www.kaggle.com/horsek\" target=\"_blank\">@horsek</a> Thanks and congrats your silver medal!<br>\nI understand your approach. Maybe it almost same effect as ours if you set the appropriate timediff-threshold(=10s in our case).</p>\n<pre><code>t_l: wifi last seen timestamp\nt_g: wifi group timestamp\nt_w: waypoint timestamp\n\nour timediff\n = t_w - t_l\n\nyour timediff\n= t_g - t_l \n= (t_g - t_l) + (t_w - t_w) \n= (t_w - t_l) + (t_g - t_w) \n= (our timediff) + (t_g - t_w)\n</code></pre>\n<p>where (t_g - t_w) should be constant if we and you use kouki's dataset.</p>\n<p>If there are any mistakes, please point them out.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1313720,
          "author_name": "horsek",
          "author_url": "",
          "post_date": "05/18/2021 17:14:59",
          "content": "<p>Thank you! Congrats to you too on your silver medal.<br>\nI see. I agree that we are essentially doing the same thing. And the fact that both of our models have shown pretty good improvement tells that you are right.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1323375,
      "author_name": "yashsharmabharatpur",
      "author_url": "",
      "post_date": "05/26/2021 07:23:23",
      "content": "<p>Thanks for Sharing this to us. <br>\nUpvote for you. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1312381": "Congratulations to all the winners, and thanks so much for hosting such an interesting competition! I'll share our solution. \n\n## Summary\n![Indoor-solution (1)](https://user-images.githubusercontent.com/43205304/118599610-c62cff00-b7ea-11eb-9040-d49651249bbe.png)\n\n\n---\n\n## dataset\nOur dataset is based on [indoor-unified-wifi-ds](https://www.kaggle.com/kokitanisaka/indoorunifiedwifids). The major differences are as follows.\n- Remove the data, if the difference between timestamp of waypoint and last seen timestamp of wifi is more than 10s. **(This is very important)**\n- The test data was also created as a wifi-based dataset in the same way as the train in order to average the timestamp during inference. So there are about three times as many lines as the original.\n- The hidden waypoints are complemented by [Linear Interpolation](https://www.kaggle.com/kuto0633/linear-interpolation-for-waypoint-in-wifi-dataset) and [Kalman Filter](https://www.kaggle.com/arnaudcapitaine/get-indoor-location-by-means-of-kalman-smoother). This can be expected to have a padding effect on the data.\n<img width=\"320\" alt=\"スクリーンショット 2021-05-18 8 41 25\" src=\"https://user-images.githubusercontent.com/43205304/118569421-e2ae4480-b7b4-11eb-9b6f-713566812365.png\">\n\n---\n## model\n### model for xy\nOur Model is based on [LSTM by Keras with Unified Wi-Fi Feats](https://www.kaggle.com/kokitanisaka/lstm-by-keras-with-unified-wi-fi-feats). Since this LSTM model is not time series based, we also use the MLP model. \nUsing features is following.\n- 80 pieces BSSID\n- 80 pieces RSSI\n- site id \n- floor (In inference, we use values which other model predict)\n\n###  model for floor\n- lightGBM (also try BiLSTM)\n- Learning as a classification\n- Create model in each site\n- Create dataset each path file as one line\n- StratifiedGroupKFold\n- features\n    - BSSID  \n    - Mean RSSI\n    - Max RSSI\n    - Extract BSSIDs that are present in both test and train.\n\n\n---\n## training (2stage) \n\n### 1st stage\n- Almost as shown in the figure above.\n- Custom loss\nEach data in our dataset has timediff(time difference between timestamp of waypoint and timestamp of wifi group).  The larger this timediff is, the larger the discrepancy between the given target and the ground truth. So we apply Weighted-Loss according to timediff.\n\n```python\n# timediff -> weight\ntimediff = df['timediff'].astype(np.float32).abs().values\nweight = 1- (timediff/np.max(timediff)) \n\nclass WeightedMSELoss(nn.Module):\n    def __init__(self):\n        super().__init__()\n        self.loss = nn.MSELoss(reduction='none')\n\n    def forward(self, input, target, weight):\n        input = input.float()\n        target = target.float()\n        weight = torch.stack((weight, weight), 1).float()  \n        loss = self.loss(input, target) * weight \n        return loss.mean()\n```\n\n### 2nd stage\nRe-learn by adding the following elements.\n- Add test data by pseudo labelig \n- Remove train data  if oof's error is over 40m (also try to remove over 20m)\n\n---\n## post processing (pp)\n\nWe use three-pp which were shared in Notebook & Discussion.\nWe repeated three-pp **6 times**. It is very effective.\n\n- [Cost minimaization](https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization)\nThere are difference between delta by sensor and delta by target. So we corrected sensor delta by the coefficients which obtained by linear regression  in each site and each floor.\n\n| | |\n| --- | --- |\n| <img width=\"362\" alt=\"スクリーンショット 2021-05-18 8 49 51\" src=\"https://user-images.githubusercontent.com/43205304/118569982-16d63500-b7b6-11eb-8f70-35f1c2c10c1c.png\"> | <img width=\"361\" alt=\"スクリーンショット 2021-05-18 8 49 59\" src=\"https://user-images.githubusercontent.com/43205304/118569988-189ff880-b7b6-11eb-9e0e-060ece210127.png\">  |\n  \n  \n- [Snap to grid](https://www.kaggle.com/robikscube/indoor-navigation-snap-to-grid-post-processing)\nWe automatically generated 3 pattern extra grids instead of hand-labeling grids like following. \n  \n| sparse | dense | edge |\n| --- | --- | --- |\n| <img width=\"495\" alt=\"スクリーンショット 2021-05-18 8 13 53\" src=\"https://user-images.githubusercontent.com/43205304/118567646-1dae7900-b7b1-11eb-9425-ffb414c7f923.png\"> | <img width=\"495\" alt=\"スクリーンショット 2021-05-18 8 14 27\" src=\"https://user-images.githubusercontent.com/43205304/118567717-42a2ec00-b7b1-11eb-957f-b18649ea3f0d.png\"> | <img width=\"492\" alt=\"スクリーンショット 2021-05-18 8 14 07\" src=\"https://user-images.githubusercontent.com/43205304/118567652-21420000-b7b1-11eb-92c3-29329d2d1fd5.png\">\n   \n   \n- [device id leak](https://www.kaggle.com/iwatatakuya/use-leakage-considering-device-id-postprocess)\n\n## ensemble\nWe did stacking and extra grid ensemble. We did post processing with 4 different patterns. Then, did ensemble by weighted average.\n\n4 pattern is here.\n① snap to grid’s threshold=None / sparse extra grid\n② snap to grid’s threshold=None / dense  extra grid\n③ snap to grid’s threshold=None / edge  extra grid\n④ snap to grid’s threshold=5 / only train grid\n\n---\n\nThank you!",
    "1312429": "Thank you for publishing a great solution.Can you tell me about the post-processing repeat?\n\nI have tried repeats and found them to be effective in some cases and vice versa. I wasn't confident about the effectiveness of repeat so I didn't use it in my final sub.\n\nWhat was your score before post-processing?\n Also, how effective were the repeats without extra grids?\n\nPersonally, I think that if there are no extra grid points, there will be no effect of repeating the process.",
    "1312465": "dehokanta Thanks for your comment.\n\n> What was your score before post-processing?\nAlso, how effective were the repeats without extra grids?\n\nOur single model's CV is here. \nThis score is used extra grid, but the one which is not used extra grid is almost same improvement.\nAfter the 3 times, the effect of pp slowed down but the CV increased, so we decided 6 times.\n\n| | CV score | \n| --- | --- | \n| CV(not pp) | 6.311 |\n| CV(pp×1) | 4.672 |  \n| CV(pp×2) | 4.377 |  \n\n> Personally, I think that if there are no extra grid points, there will be no effect of repeating the process.\n\nThere are also good effect by repeating without extra-grids in our case. By doing snap to grid(N-1 times.), the predicted waypoints closer to ground truth, so the result of Cost Minimaization(N times) also changed, and the score improved, I think.",
    "1312549": "kuto0633 Thanks for the quick reply.\nThe result of our PP (without hand label) loop was this.\n|  | Public Score |\n| --- | --- |\n|  without PP | 5.096 |\n|  PP×1 | 3.569 |\n|  PP×4 | 3.766 |\n\nI didn't use any additional grids, so Snap to Grid moved  from the correct position to the wrong position, which may be why our score didn't improve.\n\nThank you very much for valuable comments.",
    "1313046": "Thank you for sharing this great solution!\nI learned a lot from this 😃 (eg. Custom loss, Removing big timediff samples or large loss samples, pseudo labeling)\n\nI have one question about your model.\n> Since this LSTM model is not time series based, we also use the MLP model.\n\nDoes this mean you created two different models (not time-series lstm and mlp) and the latter MLP model used time-series information like gru or lstm ?\n\nThanks in advance!",
    "1313076": "yutoshibata Thanks! I'm glad such your comments.\n> you created two different models (not time-series lstm and mlp) \n\nYes, I and another teammate created LSTM, and the other created MLP.\n\n> the latter MLP model used time-series information like gru or lstm ?\n\nNo,  I mean that our LSTM is almost the same as MLP because it doesn't deal with time series. So our any model are not time-series.",
    "1313315": "kuto0633 Thanks ! I totally understand your comment 😃\nI used MLP too but CV score was 7.3....\nI think removing some samples or using KF or custom loss made this big difference.\nAnyway, thank you for explaining your solution kindly🙌",
    "1313406": "Thank you for sharing!\nI had a lot to learn from your great solutions.\n> Remove the data, if the difference between timestamp of waypoint and last seen timestamp of wifi is more than 10s. (This is very important)  \n\nActually, I was able to improve my score significantly by adopting this idea as well.\nHowever, my approach to discarding data is a little different from yours. I did not give any superiority to the importance based on the time difference from the waypoint, so I focused on the relation between last seen timestamp and wifi timestamp at the same row in the text file. For example, \n  \ntimestamp=1000, bssid=AAA, lastseen_ts=950 <-- keep\ntimestamp=1200, bssid=AAA, lastseen_ts=950 <-- discard because this information is older than the previous row\n  \nNow that I know how to skillfully utilize the time difference from the waypoint, I feel that your method is better.\nI would be happy to hear your frank opinion.",
    "1313548": "horsek Thanks and congrats your silver medal!\nI understand your approach. Maybe it almost same effect as ours if you set the appropriate timediff-threshold(=10s in our case).\n\n```\nt_l: wifi last seen timestamp\nt_g: wifi group timestamp\nt_w: waypoint timestamp\n\nour timediff\n = t_w - t_l\n\nyour timediff\n= t_g - t_l \n= (t_g - t_l) + (t_w - t_w) \n= (t_w - t_l) + (t_g - t_w) \n= (our timediff) + (t_g - t_w)\n\n```\nwhere (t_g - t_w) should be constant if we and you use kouki's dataset.\n\nIf there are any mistakes, please point them out.",
    "1313720": "Thank you! Congrats to you too on your silver medal.\nI see. I agree that we are essentially doing the same thing. And the fact that both of our models have shown pretty good improvement tells that you are right.",
    "1323375": "Thanks for Sharing this to us. \nUpvote for you."
  },
  "source": "meta"
}