{
  "id": 240176,
  "title": "1st Place Solution - Track me if you can",
  "url": "/competitions/indoor-location-navigation/writeups/track-me-if-you-can-1st-place-solution-track-me-if",
  "author_name": "",
  "post_date": "2021-05-25T18:43:40.803Z",
  "votes": 151,
  "comment_count": 18,
  "views": 0,
  "content": "<p>We are excited to share our winning solution. It is extremely unusual to win on Kaggle with such a large margin without relying on obscure leaks. Our final solution was inspired by the tremendous sharing on the forum (especially <a href=\"https://www.kaggle.com/robikscube/indoor-navigation-snap-to-grid-post-processing\" target=\"_blank\">snap  to grid</a> and <a href=\"https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization\" target=\"_blank\">cost minimzation</a> were key insights). We combined the public ideas with unique modeling and optimization insights. The dataset in this challenge was very rich and allowed for impressive progress by many of the top teams until the last day, which made this my favorite competition so far.</p>\n<p>We would also like to express our admiration for the team of <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a>. Your progress was very impressive throughout the challenge and your generous sharing on the forum inspired many of us to keep going. Your <a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/235328#1288198\" target=\"_blank\">prophecy</a> (\"To be honest, I'm afraid Tom &amp; dott team, who scores 6.2, is hiding a score now and show 0.x on the last day.\") sounded like an untenable dream at the time, but after we got all the pieces together, we started to believe in it ourselves and we are proud that we actually lived up to your expectations.</p>\n<h2>Summary of our approach</h2>\n<p>Like all top teams, we realized that this is essentially an optimization problem where you need to incorporate the relevant data modalities (predominantly WiFi + sensor data) at the trajectory level to achieve the best possible score. The key challenge was how to combine the predictions of the relevant data modalities with the discreteness of the prediction problem (about 85% of test predictions occur on X-Y locations seen in training).</p>\n<p>Most top teams took the route of iteratively interleaving continuous optimization with discretization by snapping to the grid. We went for all-out discrete optimization instead, where we only considered the training waypoints, as well as the procedurally generated likely additional waypoints (see “Waypoint generation”) for each prediction. Discretizing the optimization allowed us to implement a customized Beam Search procedure which combined the penalties of all relevant predictions and allowed us to efficiently find the best match at the trajectory level.</p>\n<p>We maintained a fixed holdout set (instead of cross validation), which was achieved by randomly selecting from the longer training trajectories. Our validation score was very well correlated with the private leaderboard score.</p>\n<p>In what follows, we intend to focus on the insights that were not shared before on the forum.</p>\n<h2>Base models</h2>\n<h3>WiFi data</h3>\n<p>We considered 3 modeling approaches for using WiFi data: NN, LightGBM and K-Nearest Neighbors. But only LightGBM and K-Nearest Neighbours were used in the final pipeline, having the highest accuracy. For both of them we assume that the floor is given.</p>\n<p>The LightGBM models predict X and Y coordinates at the time of a WiFi observation. There are 100+ models, one per each site * level. A particular boost in performance was observed by taking the maximum signal strength in a window of +- 8s each observation. Additionally magnetic and bluetooth data was used, but didn’t give a significant boost, resulting in a distance error of 7.25.</p>\n<p>For the kNN model, all we had to do was to find a proper distance function between two WiFi observations, and compare WiFi observations at inference time with all WiFi observations seen during training. We settled on a distance function where you combine the average rssid strength difference for the shared devices with the fraction of shared devices. Less weight is given to device observations that are delayed (time difference between WiFi t1 and t2). A pointwise weighted kNN prediction got us to a distance error of 5.8</p>\n<p>The main benefit of kNN over LightGBM is that it allows for a highly nonlinear penalty function which can incorporate more information about the layout of the floor without risking overfitting (we only have a handful of parameters in the distance function which are shared between all sites and floors).</p>\n<h3>Sensor data</h3>\n<p>Sensor data was clearly the key information to reconstruct the trajectories in this competition. When using sensor data, we focused on predicting segments - pieces of trajectories from one waypoint to the next one. We ended up using 3 different sensor data models with different targets.</p>\n<p>The first sensor model is targeted at predicting the relative movement of a segment, i.e. changes in absolute coordinates between two consequent waypoints. The model uses 9 inputs (acce, gyro and ahrs of all 3 axes) and has a relatively simple structure: Conv1d + GRU + Conv1d + head. The presence of GRU is unnecessary, similar results can be achieved when GRU is replaced with Conv1d, kernels sizes are set wider and dilations are added. The model has MAE = 1.04 on validation (averaged over the X and Y dimensions).</p>\n<p>After analyzing the errors of the model, we saw that the predicted angle of the movement direction is far from perfect, though never off by more than pi/2 radians. That can indicate that the device was not properly calibrated.</p>\n<p><a href=\"https://www.kaggle.com/c/indoor-location-navigation/leaderboard\" target=\"_blank\"><img src=\"https://i.ibb.co/2hNvxfP/sensor-error.png\"></a></p>\n<p>Even when the direction was not properly calibrated, the sensor data still can describe the movement well, just in slightly rotated coordinate axes. To take advantage of that, we fitted the second sensor data model, where the target relative movement was expressed not in the original coordinates, but relative to the previous segment, as illustrated below:</p>\n<p><a href=\"https://www.kaggle.com/c/indoor-location-navigation/leaderboard\" target=\"_blank\"><img src=\"https://i.ibb.co/StB7bqL/rotation-invariant.png\"></a></p>\n<p>The architecture of the NN remained the same, but this time we pass 2 consecutive segments (AB and BC) to the model. The NN is applied to both segments, after that the second segment (BC) is rotated so that the prediction of the first segment (AB) is along the first axis. The rotated vector BC is the output of the model.</p>\n<p>This calibration issue also leads to a bias of the predicted distance between the visited waypoints - the relative movement model underestimates it (6.65 vs 6.91 on average in validation). The third model, fitted to predict the distance between the consecutive waypoints has a much smaller bias and, as the result, is better in predicting the traveled distance between waypoints (MAE of 0.67 vs 0.73). The RMSE of the distance model was 0.92. The architecture of the model remained the same, with only the head of the network changed accordingly.</p>\n<h2>Data leaks</h2>\n<p>We are grateful to <a href=\"https://www.kaggle.com/chris62\" target=\"_blank\">@chris62</a> for <a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/234543\" target=\"_blank\">sharing the data leaks publicly</a>. At the time of the post, we were only aware of the time leak.</p>\n<p>The device leak was very useful to identify periods in time where the sensor data was unreliable, since the errors are clearly time dependent. This enabled us to make the optimization less reliant on sensor predictions when we predict them to be noisy.</p>\n<p><a href=\"https://www.kaggle.com/c/indoor-location-navigation/leaderboard\" target=\"_blank\"><img src=\"https://i.ibb.co/cC6xkc9/sensor-angle-error.png\"></a></p>\n<p>We realized that by combining the time and device leak, one could see that apparently different device ids are likely coming from the same phone. It turns out that the time between sensor observations (after filtering outliers) correlates very well with the device ids, which could enable you to group different device ids. In the end, we never used these fused device ids, but think that there could be more to be learned from our observation.</p>\n<p><a href=\"https://www.kaggle.com/c/indoor-location-navigation/leaderboard\" target=\"_blank\"><img src=\"https://i.ibb.co/0Jbh69q/device-leak.png\"></a></p>\n<p>Our test floor predictions combined the time and device leak with WiFi distances. Sadly, we got 102 private floor predictions wrong (about 0.18 total score penalty). Trajectory ”862a4ac32755d252c6948424”, for example, should apparently be F5 instead of F4.</p>\n<p><a href=\"https://www.kaggle.com/c/indoor-location-navigation/leaderboard\" target=\"_blank\"><img src=\"https://i.ibb.co/s23wrc5/floor-prediction.png\"></a></p>\n<h2>Leaderboard probing</h2>\n<p>We quickly realized that probing was a viable strategy, since the test data set (626) was split by entire trajectories. After a couple of days of probing, we identified the 100 public test trajectories (1527 predictions in total). This was mostly useful for understanding the difference in distribution of the test data between the public and private set. The private set is clearly much harder, since it contains some notoriously hard trajectories (e.g. “e83a1c294b5d138339149fcb” and “7d401b038d6fdb08f4f8197d”). It also enabled us to speed up our submission pipeline, since we would only have to generate ~15% of the test predictions.</p>\n<h2>Waypoint generation</h2>\n<p>There is a lot more structure to the waypoints than merely filling the empty space in corridors. We built our solution to attempt to fill plausible waypoint locations without adding too much noise, especially in sensitive areas close to train waypoints.</p>\n<p>Our approach consists roughly of: </p>\n<ul>\n<li>Get a clean map of the corridors from the floor map GeoJSON</li>\n<li>Identify waypoints along walls based on the distance to the nearest wall.</li>\n<li>Gather statistics about nearest neighbor euclidean distance, distance to the wall and distance between points along the wall (project wall points to the wall line)</li>\n<li>Fill in open spots along the wall line using the distance stats and linear referencing. Corners get special consideration.</li>\n<li>Fill inner points aligned with known and generated wall points.</li>\n<li>The room layout may give overlapping generated waypoints. Resolve these to the centroid of small local clusters based on global distance stats or medium local cluster stats.</li>\n<li>Apply a hierarchy of filtering too close points: Known waypoints &gt; corner points &gt; wall points &gt; inner points</li>\n</ul>\n<p>Because it is hard to trust CV in this challenge, the parameters were tuned by hand and we mainly used our eyes to assess progress.</p>\n<p>Example of our generated waypoints:</p>\n<p><a href=\"https://www.kaggle.com/c/indoor-location-navigation/leaderboard\" target=\"_blank\"><img src=\"https://i.ibb.co/8b4NbnM/additional-grid-example.png\"></a></p>\n<h2>Optimization</h2>\n<p>The discrete optimization is at the heart of our solution. It was obvious that search for the best trajectory would have to scale linearly as a function of the trajectory length, so exhaustive search was out of the question. Beam Search to the rescue! The optimization works as follows:</p>\n<p>A) Start by considering up to 2000 most likely initial waypoints, and compute a penalty for the starting point (only WiFi in our final submission)<br>\nB) For the trajectory of length L so far, consider 100 likely candidates for the next waypoint<br>\nC) Compute a penalty for the most recent segment, for the 2000*100 considered options<br>\nD) Order the 2000*100 trajectory candidates of length L+1 by their total penalty<br>\nE) Drop the candidates that don’t belong to the top 2000 and go back to B until the trajectory is completed</p>\n<p>Our prediction is then simply the trajectory with the lowest overall penalty. During validation, we always keep the best trajectory around, so we can understand where the optimization drops the ball. After weeks of tuning the optimization, we are now confident that we will almost always select the trajectory with the lowest optimization error.<br>\nStep B discards next step waypoints that are not in the half plane of the direction of the sensor prediction. We also prefer next step waypoints that are at the approximate predicted direction and distance from the previous waypoint.</p>\n<p>In our final submissions, we consider 7 types of penalties:</p>\n<ol>\n<li><strong>WiFi</strong>: prefer waypoints where the inference WiFi signal is close to the 20 nearest train neighbors of that waypoint location.<br>\nWe also generate a small boost for the cosine similarity between the vector of the segment, and the vector of the WiFi best guess. This can be interpreted as: does the WiFi movement agree with the proposed segment direction.</li>\n<li><strong>Relative movement angle</strong>: Based on the angle between the last two segments.</li>\n<li><strong>Relative movement coordinate</strong>: Based on the independent X and Y differences between the predicted relative movement and the proposed segment. </li>\n<li><strong>Pairwise integrated relative movement coordinate</strong>. We also penalize predictions that are not consistent at the trajectory level. Every (L+1)th waypoint is assessed for compatibility with waypoints 1 through L by integrating the predicted relative movements, and comparing that integrated prediction with the vector from each waypoint to waypoint L+1.</li>\n<li><strong>Distance based</strong>: Linearly increasing penalty as you move more or less far between waypoints.</li>\n<li><strong>Time leak</strong>: Apply a fixed penalty for not agreeing with the edge points of neighboring trajectories, when those trajectories seem to be at the same location and are close in time. Additionally, apply a linearly increasing penalty for moving further away from reliable edge points.</li>\n<li><strong>Off grid penalty</strong>: In order to bias the optimization towards known grid points, we apply a penalty which increases as a function of the nearest known grid point. We also apply an additional penalty for selecting additional grid points in a region of dense grid points.</li>\n</ol>\n<p>On top of that, we disallow selecting the same waypoint in two subsequent steps. We also adjust the weight of the sensor penalties based on the uncertainty of those predictions (achieved through the time and device leak).</p>\n<p>The hyperparameters were tuned with Bayesian optimization. We spent a lot of time looking at our prime misclassifications and estimate that more than half of the error we make is due to inconsistencies in the data. </p>\n<p>Below you see a snapshot of the outcome of the optimization for a test trajectory, together with the most relevant predictions that the optimization builds on.</p>\n<p><a href=\"https://www.kaggle.com/c/indoor-location-navigation/leaderboard\" target=\"_blank\"><img src=\"https://i.ibb.co/Fbms8nT/optimization.png\"></a></p>\n<h2>Ensembling</h2>\n<p>During the last two days, we had a hard time agreeing on what additional waypoint grid to choose. If you add too many additional waypoints, the optimization can sometimes pick shifted trajectories. However, if you don’t add enough waypoints, the predictions can be drastically wrong, because of the discrete nature of the optimization.</p>\n<p>In the end, we realized that we didn’t have to choose a single grid! Our final submissions generate predictions with 3 different grids:</p>\n<ol>\n<li>Only additional wall waypoints</li>\n<li>Additional wall waypoints + sparse inner waypoints</li>\n<li>Dense additional wall waypoints + dense inner waypoints</li>\n</ol>\n<p>We select the prediction with the lowest optimization penalty, where our submissions vary in the priority corrections. Ensembling enabled us to mostly stick with the simple grid, except when it resulted in a significant drop of the optimization penalty. Both final submissions boosted our score by about 10cm, relative to only using the sparse grid. </p>\n<h2>Final thoughts</h2>\n<p>We are happy to share all of our code in <a href=\"https://github.com/ttvand/Indoor-Location-Navigation-Public\" target=\"_blank\">this public repository</a>. The repository contains all our competition code, both the used and unused bits of our final solution. We also added a main script which should generate our approximate final submissions.</p>",
  "messages": [
    {
      "id": "1313851",
      "postDate": "05/18/2021 18:40:06",
      "content": "<p>We are excited to share our winning solution. It is extremely unusual to win on Kaggle with such a large margin without relying on obscure leaks. Our final solution was inspired by the tremendous sharing on the forum (especially <a href=\"https://www.kaggle.com/robikscube/indoor-navigation-snap-to-grid-post-processing\" target=\"_blank\">snap  to grid</a> and <a href=\"https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization\" target=\"_blank\">cost minimzation</a> were key insights). We combined the public ideas with unique modeling and optimization insights. The dataset in this challenge was very rich and allowed for impressive progress by many of the top teams until the last day, which made this my favorite competition so far.</p>\n<p>We would also like to express our admiration for the team of <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a>. Your progress was very impressive throughout the challenge and your generous sharing on the forum inspired many of us to keep going. Your <a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/235328#1288198\" target=\"_blank\">prophecy</a> (\"To be honest, I'm afraid Tom &amp; dott team, who scores 6.2, is hiding a score now and show 0.x on the last day.\") sounded like an untenable dream at the time, but after we got all the pieces together, we started to believe in it ourselves and we are proud that we actually lived up to your expectations.</p>\n<h2>Summary of our approach</h2>\n<p>Like all top teams, we realized that this is essentially an optimization problem where you need to incorporate the relevant data modalities (predominantly WiFi + sensor data) at the trajectory level to achieve the best possible score. The key challenge was how to combine the predictions of the relevant data modalities with the discreteness of the prediction problem (about 85% of test predictions occur on X-Y locations seen in training).</p>\n<p>Most top teams took the route of iteratively interleaving continuous optimization with discretization by snapping to the grid. We went for all-out discrete optimization instead, where we only considered the training waypoints, as well as the procedurally generated likely additional waypoints (see “Waypoint generation”) for each prediction. Discretizing the optimization allowed us to implement a customized Beam Search procedure which combined the penalties of all relevant predictions and allowed us to efficiently find the best match at the trajectory level.</p>\n<p>We maintained a fixed holdout set (instead of cross validation), which was achieved by randomly selecting from the longer training trajectories. Our validation score was very well correlated with the private leaderboard score.</p>\n<p>In what follows, we intend to focus on the insights that were not shared before on the forum.</p>\n<h2>Base models</h2>\n<h3>WiFi data</h3>\n<p>We considered 3 modeling approaches for using WiFi data: NN, LightGBM and K-Nearest Neighbors. But only LightGBM and K-Nearest Neighbours were used in the final pipeline, having the highest accuracy. For both of them we assume that the floor is given.</p>\n<p>The LightGBM models predict X and Y coordinates at the time of a WiFi observation. There are 100+ models, one per each site * level. A particular boost in performance was observed by taking the maximum signal strength in a window of +- 8s each observation. Additionally magnetic and bluetooth data was used, but didn’t give a significant boost, resulting in a distance error of 7.25.</p>\n<p>For the kNN model, all we had to do was to find a proper distance function between two WiFi observations, and compare WiFi observations at inference time with all WiFi observations seen during training. We settled on a distance function where you combine the average rssid strength difference for the shared devices with the fraction of shared devices. Less weight is given to device observations that are delayed (time difference between WiFi t1 and t2). A pointwise weighted kNN prediction got us to a distance error of 5.8</p>\n<p>The main benefit of kNN over LightGBM is that it allows for a highly nonlinear penalty function which can incorporate more information about the layout of the floor without risking overfitting (we only have a handful of parameters in the distance function which are shared between all sites and floors).</p>\n<h3>Sensor data</h3>\n<p>Sensor data was clearly the key information to reconstruct the trajectories in this competition. When using sensor data, we focused on predicting segments - pieces of trajectories from one waypoint to the next one. We ended up using 3 different sensor data models with different targets.</p>\n<p>The first sensor model is targeted at predicting the relative movement of a segment, i.e. changes in absolute coordinates between two consequent waypoints. The model uses 9 inputs (acce, gyro and ahrs of all 3 axes) and has a relatively simple structure: Conv1d + GRU + Conv1d + head. The presence of GRU is unnecessary, similar results can be achieved when GRU is replaced with Conv1d, kernels sizes are set wider and dilations are added. The model has MAE = 1.04 on validation (averaged over the X and Y dimensions).</p>\n<p>After analyzing the errors of the model, we saw that the predicted angle of the movement direction is far from perfect, though never off by more than pi/2 radians. That can indicate that the device was not properly calibrated.</p>\n<p><a href=\"https://www.kaggle.com/c/indoor-location-navigation/leaderboard\" target=\"_blank\"><img src=\"https://i.ibb.co/2hNvxfP/sensor-error.png\"></a></p>\n<p>Even when the direction was not properly calibrated, the sensor data still can describe the movement well, just in slightly rotated coordinate axes. To take advantage of that, we fitted the second sensor data model, where the target relative movement was expressed not in the original coordinates, but relative to the previous segment, as illustrated below:</p>\n<p><a href=\"https://www.kaggle.com/c/indoor-location-navigation/leaderboard\" target=\"_blank\"><img src=\"https://i.ibb.co/StB7bqL/rotation-invariant.png\"></a></p>\n<p>The architecture of the NN remained the same, but this time we pass 2 consecutive segments (AB and BC) to the model. The NN is applied to both segments, after that the second segment (BC) is rotated so that the prediction of the first segment (AB) is along the first axis. The rotated vector BC is the output of the model.</p>\n<p>This calibration issue also leads to a bias of the predicted distance between the visited waypoints - the relative movement model underestimates it (6.65 vs 6.91 on average in validation). The third model, fitted to predict the distance between the consecutive waypoints has a much smaller bias and, as the result, is better in predicting the traveled distance between waypoints (MAE of 0.67 vs 0.73). The RMSE of the distance model was 0.92. The architecture of the model remained the same, with only the head of the network changed accordingly.</p>\n<h2>Data leaks</h2>\n<p>We are grateful to <a href=\"https://www.kaggle.com/chris62\" target=\"_blank\">@chris62</a> for <a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/234543\" target=\"_blank\">sharing the data leaks publicly</a>. At the time of the post, we were only aware of the time leak.</p>\n<p>The device leak was very useful to identify periods in time where the sensor data was unreliable, since the errors are clearly time dependent. This enabled us to make the optimization less reliant on sensor predictions when we predict them to be noisy.</p>\n<p><a href=\"https://www.kaggle.com/c/indoor-location-navigation/leaderboard\" target=\"_blank\"><img src=\"https://i.ibb.co/cC6xkc9/sensor-angle-error.png\"></a></p>\n<p>We realized that by combining the time and device leak, one could see that apparently different device ids are likely coming from the same phone. It turns out that the time between sensor observations (after filtering outliers) correlates very well with the device ids, which could enable you to group different device ids. In the end, we never used these fused device ids, but think that there could be more to be learned from our observation.</p>\n<p><a href=\"https://www.kaggle.com/c/indoor-location-navigation/leaderboard\" target=\"_blank\"><img src=\"https://i.ibb.co/0Jbh69q/device-leak.png\"></a></p>\n<p>Our test floor predictions combined the time and device leak with WiFi distances. Sadly, we got 102 private floor predictions wrong (about 0.18 total score penalty). Trajectory ”862a4ac32755d252c6948424”, for example, should apparently be F5 instead of F4.</p>\n<p><a href=\"https://www.kaggle.com/c/indoor-location-navigation/leaderboard\" target=\"_blank\"><img src=\"https://i.ibb.co/s23wrc5/floor-prediction.png\"></a></p>\n<h2>Leaderboard probing</h2>\n<p>We quickly realized that probing was a viable strategy, since the test data set (626) was split by entire trajectories. After a couple of days of probing, we identified the 100 public test trajectories (1527 predictions in total). This was mostly useful for understanding the difference in distribution of the test data between the public and private set. The private set is clearly much harder, since it contains some notoriously hard trajectories (e.g. “e83a1c294b5d138339149fcb” and “7d401b038d6fdb08f4f8197d”). It also enabled us to speed up our submission pipeline, since we would only have to generate ~15% of the test predictions.</p>\n<h2>Waypoint generation</h2>\n<p>There is a lot more structure to the waypoints than merely filling the empty space in corridors. We built our solution to attempt to fill plausible waypoint locations without adding too much noise, especially in sensitive areas close to train waypoints.</p>\n<p>Our approach consists roughly of: </p>\n<ul>\n<li>Get a clean map of the corridors from the floor map GeoJSON</li>\n<li>Identify waypoints along walls based on the distance to the nearest wall.</li>\n<li>Gather statistics about nearest neighbor euclidean distance, distance to the wall and distance between points along the wall (project wall points to the wall line)</li>\n<li>Fill in open spots along the wall line using the distance stats and linear referencing. Corners get special consideration.</li>\n<li>Fill inner points aligned with known and generated wall points.</li>\n<li>The room layout may give overlapping generated waypoints. Resolve these to the centroid of small local clusters based on global distance stats or medium local cluster stats.</li>\n<li>Apply a hierarchy of filtering too close points: Known waypoints &gt; corner points &gt; wall points &gt; inner points</li>\n</ul>\n<p>Because it is hard to trust CV in this challenge, the parameters were tuned by hand and we mainly used our eyes to assess progress.</p>\n<p>Example of our generated waypoints:</p>\n<p><a href=\"https://www.kaggle.com/c/indoor-location-navigation/leaderboard\" target=\"_blank\"><img src=\"https://i.ibb.co/8b4NbnM/additional-grid-example.png\"></a></p>\n<h2>Optimization</h2>\n<p>The discrete optimization is at the heart of our solution. It was obvious that search for the best trajectory would have to scale linearly as a function of the trajectory length, so exhaustive search was out of the question. Beam Search to the rescue! The optimization works as follows:</p>\n<p>A) Start by considering up to 2000 most likely initial waypoints, and compute a penalty for the starting point (only WiFi in our final submission)<br>\nB) For the trajectory of length L so far, consider 100 likely candidates for the next waypoint<br>\nC) Compute a penalty for the most recent segment, for the 2000*100 considered options<br>\nD) Order the 2000*100 trajectory candidates of length L+1 by their total penalty<br>\nE) Drop the candidates that don’t belong to the top 2000 and go back to B until the trajectory is completed</p>\n<p>Our prediction is then simply the trajectory with the lowest overall penalty. During validation, we always keep the best trajectory around, so we can understand where the optimization drops the ball. After weeks of tuning the optimization, we are now confident that we will almost always select the trajectory with the lowest optimization error.<br>\nStep B discards next step waypoints that are not in the half plane of the direction of the sensor prediction. We also prefer next step waypoints that are at the approximate predicted direction and distance from the previous waypoint.</p>\n<p>In our final submissions, we consider 7 types of penalties:</p>\n<ol>\n<li><strong>WiFi</strong>: prefer waypoints where the inference WiFi signal is close to the 20 nearest train neighbors of that waypoint location.<br>\nWe also generate a small boost for the cosine similarity between the vector of the segment, and the vector of the WiFi best guess. This can be interpreted as: does the WiFi movement agree with the proposed segment direction.</li>\n<li><strong>Relative movement angle</strong>: Based on the angle between the last two segments.</li>\n<li><strong>Relative movement coordinate</strong>: Based on the independent X and Y differences between the predicted relative movement and the proposed segment. </li>\n<li><strong>Pairwise integrated relative movement coordinate</strong>. We also penalize predictions that are not consistent at the trajectory level. Every (L+1)th waypoint is assessed for compatibility with waypoints 1 through L by integrating the predicted relative movements, and comparing that integrated prediction with the vector from each waypoint to waypoint L+1.</li>\n<li><strong>Distance based</strong>: Linearly increasing penalty as you move more or less far between waypoints.</li>\n<li><strong>Time leak</strong>: Apply a fixed penalty for not agreeing with the edge points of neighboring trajectories, when those trajectories seem to be at the same location and are close in time. Additionally, apply a linearly increasing penalty for moving further away from reliable edge points.</li>\n<li><strong>Off grid penalty</strong>: In order to bias the optimization towards known grid points, we apply a penalty which increases as a function of the nearest known grid point. We also apply an additional penalty for selecting additional grid points in a region of dense grid points.</li>\n</ol>\n<p>On top of that, we disallow selecting the same waypoint in two subsequent steps. We also adjust the weight of the sensor penalties based on the uncertainty of those predictions (achieved through the time and device leak).</p>\n<p>The hyperparameters were tuned with Bayesian optimization. We spent a lot of time looking at our prime misclassifications and estimate that more than half of the error we make is due to inconsistencies in the data. </p>\n<p>Below you see a snapshot of the outcome of the optimization for a test trajectory, together with the most relevant predictions that the optimization builds on.</p>\n<p><a href=\"https://www.kaggle.com/c/indoor-location-navigation/leaderboard\" target=\"_blank\"><img src=\"https://i.ibb.co/Fbms8nT/optimization.png\"></a></p>\n<h2>Ensembling</h2>\n<p>During the last two days, we had a hard time agreeing on what additional waypoint grid to choose. If you add too many additional waypoints, the optimization can sometimes pick shifted trajectories. However, if you don’t add enough waypoints, the predictions can be drastically wrong, because of the discrete nature of the optimization.</p>\n<p>In the end, we realized that we didn’t have to choose a single grid! Our final submissions generate predictions with 3 different grids:</p>\n<ol>\n<li>Only additional wall waypoints</li>\n<li>Additional wall waypoints + sparse inner waypoints</li>\n<li>Dense additional wall waypoints + dense inner waypoints</li>\n</ol>\n<p>We select the prediction with the lowest optimization penalty, where our submissions vary in the priority corrections. Ensembling enabled us to mostly stick with the simple grid, except when it resulted in a significant drop of the optimization penalty. Both final submissions boosted our score by about 10cm, relative to only using the sparse grid. </p>\n<h2>Final thoughts</h2>\n<p>We are happy to share all of our code in <a href=\"https://github.com/ttvand/Indoor-Location-Navigation-Public\" target=\"_blank\">this public repository</a>. The repository contains all our competition code, both the used and unused bits of our final solution. We also added a main script which should generate our approximate final submissions.</p>",
      "rawMarkdown": "We are excited to share our winning solution. It is extremely unusual to win on Kaggle with such a large margin without relying on obscure leaks. Our final solution was inspired by the tremendous sharing on the forum (especially [snap  to grid](https://www.kaggle.com/robikscube/indoor-navigation-snap-to-grid-post-processing) and [cost minimzation](https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization) were key insights). We combined the public ideas with unique modeling and optimization insights. The dataset in this challenge was very rich and allowed for impressive progress by many of the top teams until the last day, which made this my favorite competition so far.\n\nWe would also like to express our admiration for the team of @mamasinkgs. Your progress was very impressive throughout the challenge and your generous sharing on the forum inspired many of us to keep going. Your [prophecy](https://www.kaggle.com/c/indoor-location-navigation/discussion/235328#1288198) (\"To be honest, I'm afraid Tom & dott team, who scores 6.2, is hiding a score now and show 0.x on the last day.\") sounded like an untenable dream at the time, but after we got all the pieces together, we started to believe in it ourselves and we are proud that we actually lived up to your expectations.\n\n\n## Summary of our approach\nLike all top teams, we realized that this is essentially an optimization problem where you need to incorporate the relevant data modalities (predominantly WiFi + sensor data) at the trajectory level to achieve the best possible score. The key challenge was how to combine the predictions of the relevant data modalities with the discreteness of the prediction problem (about 85% of test predictions occur on X-Y locations seen in training).\n\nMost top teams took the route of iteratively interleaving continuous optimization with discretization by snapping to the grid. We went for all-out discrete optimization instead, where we only considered the training waypoints, as well as the procedurally generated likely additional waypoints (see “Waypoint generation”) for each prediction. Discretizing the optimization allowed us to implement a customized Beam Search procedure which combined the penalties of all relevant predictions and allowed us to efficiently find the best match at the trajectory level.\n\nWe maintained a fixed holdout set (instead of cross validation), which was achieved by randomly selecting from the longer training trajectories. Our validation score was very well correlated with the private leaderboard score.\n\nIn what follows, we intend to focus on the insights that were not shared before on the forum.\n\n\n## Base models\n### WiFi data\nWe considered 3 modeling approaches for using WiFi data: NN, LightGBM and K-Nearest Neighbors. But only LightGBM and K-Nearest Neighbours were used in the final pipeline, having the highest accuracy. For both of them we assume that the floor is given.\n\nThe LightGBM models predict X and Y coordinates at the time of a WiFi observation. There are 100+ models, one per each site \\* level. A particular boost in performance was observed by taking the maximum signal strength in a window of +- 8s each observation. Additionally magnetic and bluetooth data was used, but didn’t give a significant boost, resulting in a distance error of 7.25.\n\nFor the kNN model, all we had to do was to find a proper distance function between two WiFi observations, and compare WiFi observations at inference time with all WiFi observations seen during training. We settled on a distance function where you combine the average rssid strength difference for the shared devices with the fraction of shared devices. Less weight is given to device observations that are delayed (time difference between WiFi t1 and t2). A pointwise weighted kNN prediction got us to a distance error of 5.8\n\nThe main benefit of kNN over LightGBM is that it allows for a highly nonlinear penalty function which can incorporate more information about the layout of the floor without risking overfitting (we only have a handful of parameters in the distance function which are shared between all sites and floors).\n\n\n### Sensor data\nSensor data was clearly the key information to reconstruct the trajectories in this competition. When using sensor data, we focused on predicting segments - pieces of trajectories from one waypoint to the next one. We ended up using 3 different sensor data models with different targets.\n\nThe first sensor model is targeted at predicting the relative movement of a segment, i.e. changes in absolute coordinates between two consequent waypoints. The model uses 9 inputs (acce, gyro and ahrs of all 3 axes) and has a relatively simple structure: Conv1d + GRU + Conv1d + head. The presence of GRU is unnecessary, similar results can be achieved when GRU is replaced with Conv1d, kernels sizes are set wider and dilations are added. The model has MAE = 1.04 on validation (averaged over the X and Y dimensions).\n\nAfter analyzing the errors of the model, we saw that the predicted angle of the movement direction is far from perfect, though never off by more than pi/2 radians. That can indicate that the device was not properly calibrated.\n\n[<img src=\"https://i.ibb.co/2hNvxfP/sensor-error.png\">](https://www.kaggle.com/c/indoor-location-navigation/leaderboard)\n\nEven when the direction was not properly calibrated, the sensor data still can describe the movement well, just in slightly rotated coordinate axes. To take advantage of that, we fitted the second sensor data model, where the target relative movement was expressed not in the original coordinates, but relative to the previous segment, as illustrated below:\n\n[<img src=\"https://i.ibb.co/StB7bqL/rotation-invariant.png\">](https://www.kaggle.com/c/indoor-location-navigation/leaderboard)\n\nThe architecture of the NN remained the same, but this time we pass 2 consecutive segments (AB and BC) to the model. The NN is applied to both segments, after that the second segment (BC) is rotated so that the prediction of the first segment (AB) is along the first axis. The rotated vector BC is the output of the model.\n\nThis calibration issue also leads to a bias of the predicted distance between the visited waypoints - the relative movement model underestimates it (6.65 vs 6.91 on average in validation). The third model, fitted to predict the distance between the consecutive waypoints has a much smaller bias and, as the result, is better in predicting the traveled distance between waypoints (MAE of 0.67 vs 0.73). The RMSE of the distance model was 0.92. The architecture of the model remained the same, with only the head of the network changed accordingly.\n\n\n## Data leaks\nWe are grateful to @chris62 for [sharing the data leaks publicly](https://www.kaggle.com/c/indoor-location-navigation/discussion/234543). At the time of the post, we were only aware of the time leak.\n\nThe device leak was very useful to identify periods in time where the sensor data was unreliable, since the errors are clearly time dependent. This enabled us to make the optimization less reliant on sensor predictions when we predict them to be noisy.\n\n[<img src=\"https://i.ibb.co/cC6xkc9/sensor-angle-error.png\">](https://www.kaggle.com/c/indoor-location-navigation/leaderboard)\n\nWe realized that by combining the time and device leak, one could see that apparently different device ids are likely coming from the same phone. It turns out that the time between sensor observations (after filtering outliers) correlates very well with the device ids, which could enable you to group different device ids. In the end, we never used these fused device ids, but think that there could be more to be learned from our observation.\n\n[<img src=\"https://i.ibb.co/0Jbh69q/device-leak.png\">](https://www.kaggle.com/c/indoor-location-navigation/leaderboard)\n\nOur test floor predictions combined the time and device leak with WiFi distances. Sadly, we got 102 private floor predictions wrong (about 0.18 total score penalty). Trajectory ”862a4ac32755d252c6948424”, for example, should apparently be F5 instead of F4.\n\n[<img src=\"https://i.ibb.co/s23wrc5/floor-prediction.png\">](https://www.kaggle.com/c/indoor-location-navigation/leaderboard)\n\n\n## Leaderboard probing\nWe quickly realized that probing was a viable strategy, since the test data set (626) was split by entire trajectories. After a couple of days of probing, we identified the 100 public test trajectories (1527 predictions in total). This was mostly useful for understanding the difference in distribution of the test data between the public and private set. The private set is clearly much harder, since it contains some notoriously hard trajectories (e.g. “e83a1c294b5d138339149fcb” and “7d401b038d6fdb08f4f8197d”). It also enabled us to speed up our submission pipeline, since we would only have to generate ~15% of the test predictions.\n\n\n## Waypoint generation\nThere is a lot more structure to the waypoints than merely filling the empty space in corridors. We built our solution to attempt to fill plausible waypoint locations without adding too much noise, especially in sensitive areas close to train waypoints.\n\nOur approach consists roughly of: \n- Get a clean map of the corridors from the floor map GeoJSON\n- Identify waypoints along walls based on the distance to the nearest wall.\n- Gather statistics about nearest neighbor euclidean distance, distance to the wall and distance between points along the wall (project wall points to the wall line)\n- Fill in open spots along the wall line using the distance stats and linear referencing. Corners get special consideration.\n- Fill inner points aligned with known and generated wall points.\n- The room layout may give overlapping generated waypoints. Resolve these to the centroid of small local clusters based on global distance stats or medium local cluster stats.\n- Apply a hierarchy of filtering too close points: Known waypoints > corner points > wall points > inner points\n\nBecause it is hard to trust CV in this challenge, the parameters were tuned by hand and we mainly used our eyes to assess progress.\n\nExample of our generated waypoints:\n\n[<img src=\"https://i.ibb.co/8b4NbnM/additional-grid-example.png\">](https://www.kaggle.com/c/indoor-location-navigation/leaderboard)\n\n\n## Optimization\nThe discrete optimization is at the heart of our solution. It was obvious that search for the best trajectory would have to scale linearly as a function of the trajectory length, so exhaustive search was out of the question. Beam Search to the rescue! The optimization works as follows:\n\n\nA) Start by considering up to 2000 most likely initial waypoints, and compute a penalty for the starting point (only WiFi in our final submission)\nB) For the trajectory of length L so far, consider 100 likely candidates for the next waypoint\nC) Compute a penalty for the most recent segment, for the 2000\\*100 considered options\nD) Order the 2000\\*100 trajectory candidates of length L+1 by their total penalty\nE) Drop the candidates that don’t belong to the top 2000 and go back to B until the trajectory is completed\n\nOur prediction is then simply the trajectory with the lowest overall penalty. During validation, we always keep the best trajectory around, so we can understand where the optimization drops the ball. After weeks of tuning the optimization, we are now confident that we will almost always select the trajectory with the lowest optimization error.\nStep B discards next step waypoints that are not in the half plane of the direction of the sensor prediction. We also prefer next step waypoints that are at the approximate predicted direction and distance from the previous waypoint.\n\nIn our final submissions, we consider 7 types of penalties:\n1. **WiFi**: prefer waypoints where the inference WiFi signal is close to the 20 nearest train neighbors of that waypoint location.\nWe also generate a small boost for the cosine similarity between the vector of the segment, and the vector of the WiFi best guess. This can be interpreted as: does the WiFi movement agree with the proposed segment direction.\n2. **Relative movement angle**: Based on the angle between the last two segments.\n3. **Relative movement coordinate**: Based on the independent X and Y differences between the predicted relative movement and the proposed segment. \n4. **Pairwise integrated relative movement coordinate**. We also penalize predictions that are not consistent at the trajectory level. Every (L+1)th waypoint is assessed for compatibility with waypoints 1 through L by integrating the predicted relative movements, and comparing that integrated prediction with the vector from each waypoint to waypoint L+1.\n5. **Distance based**: Linearly increasing penalty as you move more or less far between waypoints.\n6. **Time leak**: Apply a fixed penalty for not agreeing with the edge points of neighboring trajectories, when those trajectories seem to be at the same location and are close in time. Additionally, apply a linearly increasing penalty for moving further away from reliable edge points.\n7. **Off grid penalty**: In order to bias the optimization towards known grid points, we apply a penalty which increases as a function of the nearest known grid point. We also apply an additional penalty for selecting additional grid points in a region of dense grid points.\n\nOn top of that, we disallow selecting the same waypoint in two subsequent steps. We also adjust the weight of the sensor penalties based on the uncertainty of those predictions (achieved through the time and device leak).\n\nThe hyperparameters were tuned with Bayesian optimization. We spent a lot of time looking at our prime misclassifications and estimate that more than half of the error we make is due to inconsistencies in the data. \n\nBelow you see a snapshot of the outcome of the optimization for a test trajectory, together with the most relevant predictions that the optimization builds on.\n\n[<img src=\"https://i.ibb.co/Fbms8nT/optimization.png\">](https://www.kaggle.com/c/indoor-location-navigation/leaderboard)\n\n\n## Ensembling\nDuring the last two days, we had a hard time agreeing on what additional waypoint grid to choose. If you add too many additional waypoints, the optimization can sometimes pick shifted trajectories. However, if you don’t add enough waypoints, the predictions can be drastically wrong, because of the discrete nature of the optimization.\n\nIn the end, we realized that we didn’t have to choose a single grid! Our final submissions generate predictions with 3 different grids:\n1. Only additional wall waypoints\n2. Additional wall waypoints + sparse inner waypoints\n3. Dense additional wall waypoints + dense inner waypoints\n\nWe select the prediction with the lowest optimization penalty, where our submissions vary in the priority corrections. Ensembling enabled us to mostly stick with the simple grid, except when it resulted in a significant drop of the optimization penalty. Both final submissions boosted our score by about 10cm, relative to only using the sparse grid. \n\n\n## Final thoughts\nWe are happy to share all of our code in [this public repository](https://github.com/ttvand/Indoor-Location-Navigation-Public). The repository contains all our competition code, both the used and unused bits of our final solution. We also added a main script which should generate our approximate final submissions.",
      "votes": null
    },
    {
      "id": "1313900",
      "postDate": "05/18/2021 19:27:13",
      "content": "<p>Congratulations! Amazing solutions. I'm wondering when you reached this conclusion:</p>\n<pre><code>about 90% of test predictions occur on X-Y locations seen in training\n</code></pre>\n<p>Public LB is 15% data anyway. Or you mean  90% of test predictions occur on X-Y locations seen in training <strong>and generated waypoints</strong>?</p>",
      "rawMarkdown": "Congratulations! Amazing solutions. I'm wondering when you reached this conclusion:\n```\nabout 90% of test predictions occur on X-Y locations seen in training\n```\nPublic LB is 15% data anyway. Or you mean  90% of test predictions occur on X-Y locations seen in training **and generated waypoints**?",
      "votes": null
    },
    {
      "id": "1313913",
      "postDate": "05/18/2021 19:41:25",
      "content": "<p>About 85% of our optimisation outcomes are waypoints from the train grid. We make the leap here that the optimization is unbiased. This is likely since we observed no bias in validation.</p>",
      "rawMarkdown": "About 85% of our optimisation outcomes are waypoints from the train grid. We make the leap here that the optimization is unbiased. This is likely since we observed no bias in validation.",
      "votes": null
    },
    {
      "id": "1313983",
      "postDate": "05/18/2021 21:03:23",
      "content": "<p>Thanks Tom for sharing the solution. I strongly believe this is one of the greatest solution I saw in Kaggle. Your overall strategy is much better but similar to our strategy, because the heart of our solution is also the discrete optimization and the penalties you use are somewhat similar to the ones we use. </p>\n<p>The reason we lost seems (This is my poem, maybe I should write it in my solution😂, which I started writing just now.):  <br>\n1) The optimization technique (that is used in heuristic contests) is much better than ours. You did beam search, which is much better than the simple greedy algorithm used in our postprocessing. I tried optimization chunk by chunk. <br>\n2) Your grid generation strategy is great. I think probably we would have tried a similar strategy if you didn't hide LB score, but anyway we would have lost because of 1) and 3). <br>\n3) Our absolute position prediction is much better than yours, but MAE of my delta prediction model is about 1.05, which is much worse than yours and it should be more important. </p>\n<p>BTW, I have some questions about your waypoint generation techniques. Did you do per-site-floor parameter tuning by hand in waypoint generation? what is the number of the parameter? I thought about the similar techniques, but I thought the number of the parameter will become too big and seems I have to tune the parameters by site-floor (and I was a bit afraid it may be regarded as hand label 😭), so I gave up it.</p>\n<p>Another question is about your teammates. Which parts did your teammates (dott, areeh) focused on? I think your solution is too great and cannot be performed by solo, even if you can work 20 hours per day. </p>",
      "rawMarkdown": "Thanks Tom for sharing the solution. I strongly believe this is one of the greatest solution I saw in Kaggle. Your overall strategy is much better but similar to our strategy, because the heart of our solution is also the discrete optimization and the penalties you use are somewhat similar to the ones we use. \n\nThe reason we lost seems (This is my poem, maybe I should write it in my solution😂, which I started writing just now.):  \n1) The optimization technique (that is used in heuristic contests) is much better than ours. You did beam search, which is much better than the simple greedy algorithm used in our postprocessing. I tried optimization chunk by chunk. \n2) Your grid generation strategy is great. I think probably we would have tried a similar strategy if you didn't hide LB score, but anyway we would have lost because of 1) and 3). \n3) Our absolute position prediction is much better than yours, but MAE of my delta prediction model is about 1.05, which is much worse than yours and it should be more important. \n\nBTW, I have some questions about your waypoint generation techniques. Did you do per-site-floor parameter tuning by hand in waypoint generation? what is the number of the parameter? I thought about the similar techniques, but I thought the number of the parameter will become too big and seems I have to tune the parameters by site-floor (and I was a bit afraid it may be regarded as hand label 😭), so I gave up it.\n\nAnother question is about your teammates. Which parts did your teammates (dott, areeh) focused on? I think your solution is too great and cannot be performed by solo, even if you can work 20 hours per day.",
      "votes": null
    },
    {
      "id": "1314007",
      "postDate": "05/18/2021 21:44:03",
      "content": "<p>Your team's solutions are really interesting, thanks for sharing!</p>\n<p>I worked on the grid generation so I can answer your questions there. There are no per-site-floor parameters that are tuned by hand (this to me would be against the spirit of the competition and similar to labeling by hand). The parameters that are tuned interact with the per-site-floor statistics. For instance the parameters that change the generation the most are multipliers on the distance statistics. There are a few parameters that are really just encoding global statistics, like what is the max distance to a wall for a \"point close to the wall\" that we will ever consider. There are a lot of minor parameters that are quite easy to pick, like only merge points if they are at a ratio &lt; 1 of the intended generation distance. There are a few important parameters and many (~20+) minor ones.</p>\n<p>I wanted to design a more beautiful solution, but most of the grid generation work was a big effort the last week so I had to build more and design less to get a good result in time.</p>",
      "rawMarkdown": "Your team's solutions are really interesting, thanks for sharing!\n\nI worked on the grid generation so I can answer your questions there. There are no per-site-floor parameters that are tuned by hand (this to me would be against the spirit of the competition and similar to labeling by hand). The parameters that are tuned interact with the per-site-floor statistics. For instance the parameters that change the generation the most are multipliers on the distance statistics. There are a few parameters that are really just encoding global statistics, like what is the max distance to a wall for a \"point close to the wall\" that we will ever consider. There are a lot of minor parameters that are quite easy to pick, like only merge points if they are at a ratio < 1 of the intended generation distance. There are a few important parameters and many (~20+) minor ones.\n\nI wanted to design a more beautiful solution, but most of the grid generation work was a big effort the last week so I had to build more and design less to get a good result in time.",
      "votes": null
    },
    {
      "id": "1314019",
      "postDate": "05/18/2021 21:58:24",
      "content": "<p>areeh, congrats 1st place and thanks for your comment :) </p>\n<pre><code>There are no per-site-floor parameters that are tuned by hand.\nThe parameters that are tuned interact with the per-site-floor statistics.\n</code></pre>\n<p>Nice, it sounds very fair and really reasonable :) It's just unbelievable for me you did such a great job in only a week! </p>",
      "rawMarkdown": "areeh, congrats 1st place and thanks for your comment :) \n```\nThere are no per-site-floor parameters that are tuned by hand.\nThe parameters that are tuned interact with the per-site-floor statistics.\n```\nNice, it sounds very fair and really reasonable :) It's just unbelievable for me you did such a great job in only a week!",
      "votes": null
    },
    {
      "id": "1314027",
      "postDate": "05/18/2021 22:13:34",
      "content": "<p>Thank you. My teammates were doing excellent work so it was important to me to work hard</p>",
      "rawMarkdown": "Thank you. My teammates were doing excellent work so it was important to me to work hard",
      "votes": null
    },
    {
      "id": "1314104",
      "postDate": "05/19/2021 00:42:11",
      "content": "<p>Congrats to the entire \"Track me if you can\" <a href=\"https://www.kaggle.com/tvdwiele\" target=\"_blank\">@tvdwiele</a> <a href=\"https://www.kaggle.com/dott1718\" target=\"_blank\">@dott1718</a> and <a href=\"https://www.kaggle.com/areehdot\" target=\"_blank\">@areehdot</a>  on an amazing finish. Great solution write-up- There is a lot of unpack here and I'll need to reread a few times. I had expected the top solutions would approach this as some sort of discrete optimization problem- and your team did an amazing job creating an elegant solution. I'm glad my snap-to-grid notebook was insightful, but if I'm honest I'm sure your team would've found it quickly on your own. Bravo!</p>",
      "rawMarkdown": "Congrats to the entire \"Track me if you can\" @tvdwiele @dott1718 and @areehdot  on an amazing finish. Great solution write-up- There is a lot of unpack here and I'll need to reread a few times. I had expected the top solutions would approach this as some sort of discrete optimization problem- and your team did an amazing job creating an elegant solution. I'm glad my snap-to-grid notebook was insightful, but if I'm honest I'm sure your team would've found it quickly on your own. Bravo!",
      "votes": null
    },
    {
      "id": "1314283",
      "postDate": "05/19/2021 04:53:32",
      "content": "<p>Kudos to the team. Thank you for sharing the solution. Keep the hard-working going!!! 👍</p>",
      "rawMarkdown": "Kudos to the team. Thank you for sharing the solution. Keep the hard-working going!!! 👍",
      "votes": null
    },
    {
      "id": "1314647",
      "postDate": "05/19/2021 09:45:06",
      "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> and congrats on the 2nd place and an amazing work you and your team have done! It is particularly interesting to see that we can benchmark the accuracy of our individual models only now, and especially to see there is a large gap for sensor-based model performance. We are committed to share the entire pipeline within a week or two, so you can rerun it and check on your validation data. With regards to absolute position predictions, I was sure we can do better, but all the evidence pointed at them not being that important for the overall picture, so it was our last priority to improve those.</p>",
      "rawMarkdown": "Thanks a lot @mamasinkgs and congrats on the 2nd place and an amazing work you and your team have done! It is particularly interesting to see that we can benchmark the accuracy of our individual models only now, and especially to see there is a large gap for sensor-based model performance. We are committed to share the entire pipeline within a week or two, so you can rerun it and check on your validation data. With regards to absolute position predictions, I was sure we can do better, but all the evidence pointed at them not being that important for the overall picture, so it was our last priority to improve those.",
      "votes": null
    },
    {
      "id": "1314669",
      "postDate": "05/19/2021 10:03:42",
      "content": "<p>Generally speaking, there was no clear separation between who worked on what parts of the project. We very much functioned as a well-oiled team. I have had the pleasure to work with both Dmitry (in Kaggle) and Are (at our Norwegian startup) and we immediately hit the right collaboration dynamic. It can not be stressed enough that everyone worked very hard to be able to do as well as we did. We all worked on data exploration, brainstorming, following up on public discussions and quality checks of our predictions. The list below gives a rough approximation of the major contributors to the components discussed in our writeup:</p>\n<ul>\n<li><strong>WiFi models</strong>: Dmitry and Are tried NN modeling. Dmitry built the final LightGBM WiFi model and Tom worked on the kNN WiFi model.</li>\n<li><strong>Sensor models</strong>: Tom built some low-performant initial sensor models and Dmitry took charge of sensor modeling after that and made massive improvement over the initial models.</li>\n<li><strong>Data leaks</strong>: Mostly a collaboration between Dmitry and Tom</li>\n<li><strong>Leaderboard probing</strong>: Tom</li>\n<li><strong>Waypoint generation</strong>: Are took charge of this component and everything else related to the floor map. This was definitely one of the hardest and most important tasks in this list.</li>\n<li><strong>Optimization and ensembling</strong>: Tom</li>\n</ul>",
      "rawMarkdown": "Generally speaking, there was no clear separation between who worked on what parts of the project. We very much functioned as a well-oiled team. I have had the pleasure to work with both Dmitry (in Kaggle) and Are (at our Norwegian startup) and we immediately hit the right collaboration dynamic. It can not be stressed enough that everyone worked very hard to be able to do as well as we did. We all worked on data exploration, brainstorming, following up on public discussions and quality checks of our predictions. The list below gives a rough approximation of the major contributors to the components discussed in our writeup:\n\n- **WiFi models**: Dmitry and Are tried NN modeling. Dmitry built the final LightGBM WiFi model and Tom worked on the kNN WiFi model.\n- **Sensor models**: Tom built some low-performant initial sensor models and Dmitry took charge of sensor modeling after that and made massive improvement over the initial models.\n- **Data leaks**: Mostly a collaboration between Dmitry and Tom\n- **Leaderboard probing**: Tom\n- **Waypoint generation**: Are took charge of this component and everything else related to the floor map. This was definitely one of the hardest and most important tasks in this list.\n- **Optimization and ensembling**: Tom",
      "votes": null
    },
    {
      "id": "1315143",
      "postDate": "05/19/2021 15:40:52",
      "content": "<p>Contrats to you guys! This is one of the best solutions i saw here, i can see that a lot of analysis was required to reach up to that score. I'm looking forward to study the code once you guys release it.</p>\n<p>Thank you!</p>",
      "rawMarkdown": "Contrats to you guys! This is one of the best solutions i saw here, i can see that a lot of analysis was required to reach up to that score. I'm looking forward to study the code once you guys release it.\n\nThank you!",
      "votes": null
    },
    {
      "id": "1315161",
      "postDate": "05/19/2021 15:56:59",
      "content": "<p>Thank you for sharing the solution!</p>\n<blockquote>\n  <p>The model uses 9 inputs (acce, gyro and ahrs of all 3 axes) and has a relatively simple structure: Conv1d + GRU + Conv1d + head. </p>\n</blockquote>\n<p>Can Cond1d be thought of as a dense layer in this example? Am I correct that for a 10 seconds walk the GRU would be unrolled 10 * 50 (number of IMU recordings per sec) times?</p>",
      "rawMarkdown": "Thank you for sharing the solution!\n\n> The model uses 9 inputs (acce, gyro and ahrs of all 3 axes) and has a relatively simple structure: Conv1d + GRU + Conv1d + head. \n\nCan Cond1d be thought of as a dense layer in this example? Am I correct that for a 10 seconds walk the GRU would be unrolled 10 * 50 (number of IMU recordings per sec) times?",
      "votes": null
    },
    {
      "id": "1315263",
      "postDate": "05/19/2021 17:03:15",
      "content": "<p>A stack of 1D convolution can be thought of as a MLP which is applied to the overlapping time windows and where all weights are shared along the time dimension. <a href=\"https://www.quora.com/What-is-the-difference-between-a-convolutional-neural-network-and-a-multilayer-perceptron\" target=\"_blank\">Reference</a></p>\n<p>We would indeed unroll the GRU 500 steps in time in your example.</p>",
      "rawMarkdown": "A stack of 1D convolution can be thought of as a MLP which is applied to the overlapping time windows and where all weights are shared along the time dimension. [Reference](https://www.quora.com/What-is-the-difference-between-a-convolutional-neural-network-and-a-multilayer-perceptron)\n\nWe would indeed unroll the GRU 500 steps in time in your example.",
      "votes": null
    },
    {
      "id": "1316884",
      "postDate": "05/21/2021 01:33:44",
      "content": "<p>Congratulations to the team and thanks for sharing. Really interesting work!!</p>",
      "rawMarkdown": "Congratulations to the team and thanks for sharing. Really interesting work!!",
      "votes": null
    },
    {
      "id": "1318247",
      "postDate": "05/22/2021 06:39:49",
      "content": "<p>Thank you for sharing this!</p>",
      "rawMarkdown": "Thank you for sharing this!",
      "votes": null
    },
    {
      "id": "1322952",
      "postDate": "05/25/2021 20:20:02",
      "content": "<p>Update: We have made all our code available in a <a href=\"https://github.com/ttvand/Indoor-Location-Navigation-Public\" target=\"_blank\">public GitHub repo</a>. Enjoy!</p>",
      "rawMarkdown": "Update: We have made all our code available in a [public GitHub repo](https://github.com/ttvand/Indoor-Location-Navigation-Public). Enjoy!",
      "votes": null
    },
    {
      "id": "1715034",
      "postDate": "03/07/2022 15:24:22",
      "content": "<p>Hi there, congratulations. I am a rookie student and happen to see your work. I am interested in your sensor models. I am confused that whether the three sensor models u mentioned in the passage are corresponding to the file named 'sensor_model_movement 1 and 2' and 'sensor_model_dist' in github web page. If yes, what are the usages of files in unused folder for there are some sensor models too which are a little different. Honestly asking for teaching and being appreciated</p>",
      "rawMarkdown": "Hi there, congratulations. I am a rookie student and happen to see your work. I am interested in your sensor models. I am confused that whether the three sensor models u mentioned in the passage are corresponding to the file named 'sensor_model_movement 1 and 2' and 'sensor_model_dist' in github web page. If yes, what are the usages of files in unused folder for there are some sensor models too which are a little different. Honestly asking for teaching and being appreciated",
      "votes": null
    },
    {
      "id": "2135905",
      "postDate": "02/09/2023 00:19:06",
      "content": "<p>Thank you for sharing! 👍</p>",
      "rawMarkdown": "Thank you for sharing! 👍",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1313900,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "05/18/2021 19:27:13",
      "content": "<p>Congratulations! Amazing solutions. I'm wondering when you reached this conclusion:</p>\n<pre><code>about 90% of test predictions occur on X-Y locations seen in training\n</code></pre>\n<p>Public LB is 15% data anyway. Or you mean  90% of test predictions occur on X-Y locations seen in training <strong>and generated waypoints</strong>?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1313913,
          "author_name": "tvdwiele",
          "author_url": "",
          "post_date": "05/18/2021 19:41:25",
          "content": "<p>About 85% of our optimisation outcomes are waypoints from the train grid. We make the leap here that the optimization is unbiased. This is likely since we observed no bias in validation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1313983,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "05/18/2021 21:03:23",
      "content": "<p>Thanks Tom for sharing the solution. I strongly believe this is one of the greatest solution I saw in Kaggle. Your overall strategy is much better but similar to our strategy, because the heart of our solution is also the discrete optimization and the penalties you use are somewhat similar to the ones we use. </p>\n<p>The reason we lost seems (This is my poem, maybe I should write it in my solution😂, which I started writing just now.):  <br>\n1) The optimization technique (that is used in heuristic contests) is much better than ours. You did beam search, which is much better than the simple greedy algorithm used in our postprocessing. I tried optimization chunk by chunk. <br>\n2) Your grid generation strategy is great. I think probably we would have tried a similar strategy if you didn't hide LB score, but anyway we would have lost because of 1) and 3). <br>\n3) Our absolute position prediction is much better than yours, but MAE of my delta prediction model is about 1.05, which is much worse than yours and it should be more important. </p>\n<p>BTW, I have some questions about your waypoint generation techniques. Did you do per-site-floor parameter tuning by hand in waypoint generation? what is the number of the parameter? I thought about the similar techniques, but I thought the number of the parameter will become too big and seems I have to tune the parameters by site-floor (and I was a bit afraid it may be regarded as hand label 😭), so I gave up it.</p>\n<p>Another question is about your teammates. Which parts did your teammates (dott, areeh) focused on? I think your solution is too great and cannot be performed by solo, even if you can work 20 hours per day. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1314007,
          "author_name": "areehdot",
          "author_url": "",
          "post_date": "05/18/2021 21:44:03",
          "content": "<p>Your team's solutions are really interesting, thanks for sharing!</p>\n<p>I worked on the grid generation so I can answer your questions there. There are no per-site-floor parameters that are tuned by hand (this to me would be against the spirit of the competition and similar to labeling by hand). The parameters that are tuned interact with the per-site-floor statistics. For instance the parameters that change the generation the most are multipliers on the distance statistics. There are a few parameters that are really just encoding global statistics, like what is the max distance to a wall for a \"point close to the wall\" that we will ever consider. There are a lot of minor parameters that are quite easy to pick, like only merge points if they are at a ratio &lt; 1 of the intended generation distance. There are a few important parameters and many (~20+) minor ones.</p>\n<p>I wanted to design a more beautiful solution, but most of the grid generation work was a big effort the last week so I had to build more and design less to get a good result in time.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1314019,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "05/18/2021 21:58:24",
          "content": "<p>areeh, congrats 1st place and thanks for your comment :) </p>\n<pre><code>There are no per-site-floor parameters that are tuned by hand.\nThe parameters that are tuned interact with the per-site-floor statistics.\n</code></pre>\n<p>Nice, it sounds very fair and really reasonable :) It's just unbelievable for me you did such a great job in only a week! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1314027,
          "author_name": "areehdot",
          "author_url": "",
          "post_date": "05/18/2021 22:13:34",
          "content": "<p>Thank you. My teammates were doing excellent work so it was important to me to work hard</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1314647,
          "author_name": "dott1718",
          "author_url": "",
          "post_date": "05/19/2021 09:45:06",
          "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> and congrats on the 2nd place and an amazing work you and your team have done! It is particularly interesting to see that we can benchmark the accuracy of our individual models only now, and especially to see there is a large gap for sensor-based model performance. We are committed to share the entire pipeline within a week or two, so you can rerun it and check on your validation data. With regards to absolute position predictions, I was sure we can do better, but all the evidence pointed at them not being that important for the overall picture, so it was our last priority to improve those.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1314669,
          "author_name": "tvdwiele",
          "author_url": "",
          "post_date": "05/19/2021 10:03:42",
          "content": "<p>Generally speaking, there was no clear separation between who worked on what parts of the project. We very much functioned as a well-oiled team. I have had the pleasure to work with both Dmitry (in Kaggle) and Are (at our Norwegian startup) and we immediately hit the right collaboration dynamic. It can not be stressed enough that everyone worked very hard to be able to do as well as we did. We all worked on data exploration, brainstorming, following up on public discussions and quality checks of our predictions. The list below gives a rough approximation of the major contributors to the components discussed in our writeup:</p>\n<ul>\n<li><strong>WiFi models</strong>: Dmitry and Are tried NN modeling. Dmitry built the final LightGBM WiFi model and Tom worked on the kNN WiFi model.</li>\n<li><strong>Sensor models</strong>: Tom built some low-performant initial sensor models and Dmitry took charge of sensor modeling after that and made massive improvement over the initial models.</li>\n<li><strong>Data leaks</strong>: Mostly a collaboration between Dmitry and Tom</li>\n<li><strong>Leaderboard probing</strong>: Tom</li>\n<li><strong>Waypoint generation</strong>: Are took charge of this component and everything else related to the floor map. This was definitely one of the hardest and most important tasks in this list.</li>\n<li><strong>Optimization and ensembling</strong>: Tom</li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1314104,
      "author_name": "robikscube",
      "author_url": "",
      "post_date": "05/19/2021 00:42:11",
      "content": "<p>Congrats to the entire \"Track me if you can\" <a href=\"https://www.kaggle.com/tvdwiele\" target=\"_blank\">@tvdwiele</a> <a href=\"https://www.kaggle.com/dott1718\" target=\"_blank\">@dott1718</a> and <a href=\"https://www.kaggle.com/areehdot\" target=\"_blank\">@areehdot</a>  on an amazing finish. Great solution write-up- There is a lot of unpack here and I'll need to reread a few times. I had expected the top solutions would approach this as some sort of discrete optimization problem- and your team did an amazing job creating an elegant solution. I'm glad my snap-to-grid notebook was insightful, but if I'm honest I'm sure your team would've found it quickly on your own. Bravo!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1314283,
      "author_name": "ankitp013",
      "author_url": "",
      "post_date": "05/19/2021 04:53:32",
      "content": "<p>Kudos to the team. Thank you for sharing the solution. Keep the hard-working going!!! 👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1315143,
      "author_name": "victorasso",
      "author_url": "",
      "post_date": "05/19/2021 15:40:52",
      "content": "<p>Contrats to you guys! This is one of the best solutions i saw here, i can see that a lot of analysis was required to reach up to that score. I'm looking forward to study the code once you guys release it.</p>\n<p>Thank you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1315161,
      "author_name": "olaf2000",
      "author_url": "",
      "post_date": "05/19/2021 15:56:59",
      "content": "<p>Thank you for sharing the solution!</p>\n<blockquote>\n  <p>The model uses 9 inputs (acce, gyro and ahrs of all 3 axes) and has a relatively simple structure: Conv1d + GRU + Conv1d + head. </p>\n</blockquote>\n<p>Can Cond1d be thought of as a dense layer in this example? Am I correct that for a 10 seconds walk the GRU would be unrolled 10 * 50 (number of IMU recordings per sec) times?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1315263,
          "author_name": "tvdwiele",
          "author_url": "",
          "post_date": "05/19/2021 17:03:15",
          "content": "<p>A stack of 1D convolution can be thought of as a MLP which is applied to the overlapping time windows and where all weights are shared along the time dimension. <a href=\"https://www.quora.com/What-is-the-difference-between-a-convolutional-neural-network-and-a-multilayer-perceptron\" target=\"_blank\">Reference</a></p>\n<p>We would indeed unroll the GRU 500 steps in time in your example.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1316884,
      "author_name": "krishnagupta1",
      "author_url": "",
      "post_date": "05/21/2021 01:33:44",
      "content": "<p>Congratulations to the team and thanks for sharing. Really interesting work!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1318247,
      "author_name": "misswhite2104",
      "author_url": "",
      "post_date": "05/22/2021 06:39:49",
      "content": "<p>Thank you for sharing this!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1322952,
      "author_name": "tvdwiele",
      "author_url": "",
      "post_date": "05/25/2021 20:20:02",
      "content": "<p>Update: We have made all our code available in a <a href=\"https://github.com/ttvand/Indoor-Location-Navigation-Public\" target=\"_blank\">public GitHub repo</a>. Enjoy!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1715034,
      "author_name": "chunliangwang97",
      "author_url": "",
      "post_date": "03/07/2022 15:24:22",
      "content": "<p>Hi there, congratulations. I am a rookie student and happen to see your work. I am interested in your sensor models. I am confused that whether the three sensor models u mentioned in the passage are corresponding to the file named 'sensor_model_movement 1 and 2' and 'sensor_model_dist' in github web page. If yes, what are the usages of files in unused folder for there are some sensor models too which are a little different. Honestly asking for teaching and being appreciated</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2135905,
      "author_name": "leehour",
      "author_url": "",
      "post_date": "02/09/2023 00:19:06",
      "content": "<p>Thank you for sharing! 👍</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1313851": "We are excited to share our winning solution. It is extremely unusual to win on Kaggle with such a large margin without relying on obscure leaks. Our final solution was inspired by the tremendous sharing on the forum (especially [snap  to grid](https://www.kaggle.com/robikscube/indoor-navigation-snap-to-grid-post-processing) and [cost minimzation](https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization) were key insights). We combined the public ideas with unique modeling and optimization insights. The dataset in this challenge was very rich and allowed for impressive progress by many of the top teams until the last day, which made this my favorite competition so far.\n\nWe would also like to express our admiration for the team of @mamasinkgs. Your progress was very impressive throughout the challenge and your generous sharing on the forum inspired many of us to keep going. Your [prophecy](https://www.kaggle.com/c/indoor-location-navigation/discussion/235328#1288198) (\"To be honest, I'm afraid Tom & dott team, who scores 6.2, is hiding a score now and show 0.x on the last day.\") sounded like an untenable dream at the time, but after we got all the pieces together, we started to believe in it ourselves and we are proud that we actually lived up to your expectations.\n\n\n## Summary of our approach\nLike all top teams, we realized that this is essentially an optimization problem where you need to incorporate the relevant data modalities (predominantly WiFi + sensor data) at the trajectory level to achieve the best possible score. The key challenge was how to combine the predictions of the relevant data modalities with the discreteness of the prediction problem (about 85% of test predictions occur on X-Y locations seen in training).\n\nMost top teams took the route of iteratively interleaving continuous optimization with discretization by snapping to the grid. We went for all-out discrete optimization instead, where we only considered the training waypoints, as well as the procedurally generated likely additional waypoints (see “Waypoint generation”) for each prediction. Discretizing the optimization allowed us to implement a customized Beam Search procedure which combined the penalties of all relevant predictions and allowed us to efficiently find the best match at the trajectory level.\n\nWe maintained a fixed holdout set (instead of cross validation), which was achieved by randomly selecting from the longer training trajectories. Our validation score was very well correlated with the private leaderboard score.\n\nIn what follows, we intend to focus on the insights that were not shared before on the forum.\n\n\n## Base models\n### WiFi data\nWe considered 3 modeling approaches for using WiFi data: NN, LightGBM and K-Nearest Neighbors. But only LightGBM and K-Nearest Neighbours were used in the final pipeline, having the highest accuracy. For both of them we assume that the floor is given.\n\nThe LightGBM models predict X and Y coordinates at the time of a WiFi observation. There are 100+ models, one per each site \\* level. A particular boost in performance was observed by taking the maximum signal strength in a window of +- 8s each observation. Additionally magnetic and bluetooth data was used, but didn’t give a significant boost, resulting in a distance error of 7.25.\n\nFor the kNN model, all we had to do was to find a proper distance function between two WiFi observations, and compare WiFi observations at inference time with all WiFi observations seen during training. We settled on a distance function where you combine the average rssid strength difference for the shared devices with the fraction of shared devices. Less weight is given to device observations that are delayed (time difference between WiFi t1 and t2). A pointwise weighted kNN prediction got us to a distance error of 5.8\n\nThe main benefit of kNN over LightGBM is that it allows for a highly nonlinear penalty function which can incorporate more information about the layout of the floor without risking overfitting (we only have a handful of parameters in the distance function which are shared between all sites and floors).\n\n\n### Sensor data\nSensor data was clearly the key information to reconstruct the trajectories in this competition. When using sensor data, we focused on predicting segments - pieces of trajectories from one waypoint to the next one. We ended up using 3 different sensor data models with different targets.\n\nThe first sensor model is targeted at predicting the relative movement of a segment, i.e. changes in absolute coordinates between two consequent waypoints. The model uses 9 inputs (acce, gyro and ahrs of all 3 axes) and has a relatively simple structure: Conv1d + GRU + Conv1d + head. The presence of GRU is unnecessary, similar results can be achieved when GRU is replaced with Conv1d, kernels sizes are set wider and dilations are added. The model has MAE = 1.04 on validation (averaged over the X and Y dimensions).\n\nAfter analyzing the errors of the model, we saw that the predicted angle of the movement direction is far from perfect, though never off by more than pi/2 radians. That can indicate that the device was not properly calibrated.\n\n[<img src=\"https://i.ibb.co/2hNvxfP/sensor-error.png\">](https://www.kaggle.com/c/indoor-location-navigation/leaderboard)\n\nEven when the direction was not properly calibrated, the sensor data still can describe the movement well, just in slightly rotated coordinate axes. To take advantage of that, we fitted the second sensor data model, where the target relative movement was expressed not in the original coordinates, but relative to the previous segment, as illustrated below:\n\n[<img src=\"https://i.ibb.co/StB7bqL/rotation-invariant.png\">](https://www.kaggle.com/c/indoor-location-navigation/leaderboard)\n\nThe architecture of the NN remained the same, but this time we pass 2 consecutive segments (AB and BC) to the model. The NN is applied to both segments, after that the second segment (BC) is rotated so that the prediction of the first segment (AB) is along the first axis. The rotated vector BC is the output of the model.\n\nThis calibration issue also leads to a bias of the predicted distance between the visited waypoints - the relative movement model underestimates it (6.65 vs 6.91 on average in validation). The third model, fitted to predict the distance between the consecutive waypoints has a much smaller bias and, as the result, is better in predicting the traveled distance between waypoints (MAE of 0.67 vs 0.73). The RMSE of the distance model was 0.92. The architecture of the model remained the same, with only the head of the network changed accordingly.\n\n\n## Data leaks\nWe are grateful to @chris62 for [sharing the data leaks publicly](https://www.kaggle.com/c/indoor-location-navigation/discussion/234543). At the time of the post, we were only aware of the time leak.\n\nThe device leak was very useful to identify periods in time where the sensor data was unreliable, since the errors are clearly time dependent. This enabled us to make the optimization less reliant on sensor predictions when we predict them to be noisy.\n\n[<img src=\"https://i.ibb.co/cC6xkc9/sensor-angle-error.png\">](https://www.kaggle.com/c/indoor-location-navigation/leaderboard)\n\nWe realized that by combining the time and device leak, one could see that apparently different device ids are likely coming from the same phone. It turns out that the time between sensor observations (after filtering outliers) correlates very well with the device ids, which could enable you to group different device ids. In the end, we never used these fused device ids, but think that there could be more to be learned from our observation.\n\n[<img src=\"https://i.ibb.co/0Jbh69q/device-leak.png\">](https://www.kaggle.com/c/indoor-location-navigation/leaderboard)\n\nOur test floor predictions combined the time and device leak with WiFi distances. Sadly, we got 102 private floor predictions wrong (about 0.18 total score penalty). Trajectory ”862a4ac32755d252c6948424”, for example, should apparently be F5 instead of F4.\n\n[<img src=\"https://i.ibb.co/s23wrc5/floor-prediction.png\">](https://www.kaggle.com/c/indoor-location-navigation/leaderboard)\n\n\n## Leaderboard probing\nWe quickly realized that probing was a viable strategy, since the test data set (626) was split by entire trajectories. After a couple of days of probing, we identified the 100 public test trajectories (1527 predictions in total). This was mostly useful for understanding the difference in distribution of the test data between the public and private set. The private set is clearly much harder, since it contains some notoriously hard trajectories (e.g. “e83a1c294b5d138339149fcb” and “7d401b038d6fdb08f4f8197d”). It also enabled us to speed up our submission pipeline, since we would only have to generate ~15% of the test predictions.\n\n\n## Waypoint generation\nThere is a lot more structure to the waypoints than merely filling the empty space in corridors. We built our solution to attempt to fill plausible waypoint locations without adding too much noise, especially in sensitive areas close to train waypoints.\n\nOur approach consists roughly of: \n- Get a clean map of the corridors from the floor map GeoJSON\n- Identify waypoints along walls based on the distance to the nearest wall.\n- Gather statistics about nearest neighbor euclidean distance, distance to the wall and distance between points along the wall (project wall points to the wall line)\n- Fill in open spots along the wall line using the distance stats and linear referencing. Corners get special consideration.\n- Fill inner points aligned with known and generated wall points.\n- The room layout may give overlapping generated waypoints. Resolve these to the centroid of small local clusters based on global distance stats or medium local cluster stats.\n- Apply a hierarchy of filtering too close points: Known waypoints > corner points > wall points > inner points\n\nBecause it is hard to trust CV in this challenge, the parameters were tuned by hand and we mainly used our eyes to assess progress.\n\nExample of our generated waypoints:\n\n[<img src=\"https://i.ibb.co/8b4NbnM/additional-grid-example.png\">](https://www.kaggle.com/c/indoor-location-navigation/leaderboard)\n\n\n## Optimization\nThe discrete optimization is at the heart of our solution. It was obvious that search for the best trajectory would have to scale linearly as a function of the trajectory length, so exhaustive search was out of the question. Beam Search to the rescue! The optimization works as follows:\n\n\nA) Start by considering up to 2000 most likely initial waypoints, and compute a penalty for the starting point (only WiFi in our final submission)\nB) For the trajectory of length L so far, consider 100 likely candidates for the next waypoint\nC) Compute a penalty for the most recent segment, for the 2000\\*100 considered options\nD) Order the 2000\\*100 trajectory candidates of length L+1 by their total penalty\nE) Drop the candidates that don’t belong to the top 2000 and go back to B until the trajectory is completed\n\nOur prediction is then simply the trajectory with the lowest overall penalty. During validation, we always keep the best trajectory around, so we can understand where the optimization drops the ball. After weeks of tuning the optimization, we are now confident that we will almost always select the trajectory with the lowest optimization error.\nStep B discards next step waypoints that are not in the half plane of the direction of the sensor prediction. We also prefer next step waypoints that are at the approximate predicted direction and distance from the previous waypoint.\n\nIn our final submissions, we consider 7 types of penalties:\n1. **WiFi**: prefer waypoints where the inference WiFi signal is close to the 20 nearest train neighbors of that waypoint location.\nWe also generate a small boost for the cosine similarity between the vector of the segment, and the vector of the WiFi best guess. This can be interpreted as: does the WiFi movement agree with the proposed segment direction.\n2. **Relative movement angle**: Based on the angle between the last two segments.\n3. **Relative movement coordinate**: Based on the independent X and Y differences between the predicted relative movement and the proposed segment. \n4. **Pairwise integrated relative movement coordinate**. We also penalize predictions that are not consistent at the trajectory level. Every (L+1)th waypoint is assessed for compatibility with waypoints 1 through L by integrating the predicted relative movements, and comparing that integrated prediction with the vector from each waypoint to waypoint L+1.\n5. **Distance based**: Linearly increasing penalty as you move more or less far between waypoints.\n6. **Time leak**: Apply a fixed penalty for not agreeing with the edge points of neighboring trajectories, when those trajectories seem to be at the same location and are close in time. Additionally, apply a linearly increasing penalty for moving further away from reliable edge points.\n7. **Off grid penalty**: In order to bias the optimization towards known grid points, we apply a penalty which increases as a function of the nearest known grid point. We also apply an additional penalty for selecting additional grid points in a region of dense grid points.\n\nOn top of that, we disallow selecting the same waypoint in two subsequent steps. We also adjust the weight of the sensor penalties based on the uncertainty of those predictions (achieved through the time and device leak).\n\nThe hyperparameters were tuned with Bayesian optimization. We spent a lot of time looking at our prime misclassifications and estimate that more than half of the error we make is due to inconsistencies in the data. \n\nBelow you see a snapshot of the outcome of the optimization for a test trajectory, together with the most relevant predictions that the optimization builds on.\n\n[<img src=\"https://i.ibb.co/Fbms8nT/optimization.png\">](https://www.kaggle.com/c/indoor-location-navigation/leaderboard)\n\n\n## Ensembling\nDuring the last two days, we had a hard time agreeing on what additional waypoint grid to choose. If you add too many additional waypoints, the optimization can sometimes pick shifted trajectories. However, if you don’t add enough waypoints, the predictions can be drastically wrong, because of the discrete nature of the optimization.\n\nIn the end, we realized that we didn’t have to choose a single grid! Our final submissions generate predictions with 3 different grids:\n1. Only additional wall waypoints\n2. Additional wall waypoints + sparse inner waypoints\n3. Dense additional wall waypoints + dense inner waypoints\n\nWe select the prediction with the lowest optimization penalty, where our submissions vary in the priority corrections. Ensembling enabled us to mostly stick with the simple grid, except when it resulted in a significant drop of the optimization penalty. Both final submissions boosted our score by about 10cm, relative to only using the sparse grid. \n\n\n## Final thoughts\nWe are happy to share all of our code in [this public repository](https://github.com/ttvand/Indoor-Location-Navigation-Public). The repository contains all our competition code, both the used and unused bits of our final solution. We also added a main script which should generate our approximate final submissions.",
    "1313900": "Congratulations! Amazing solutions. I'm wondering when you reached this conclusion:\n```\nabout 90% of test predictions occur on X-Y locations seen in training\n```\nPublic LB is 15% data anyway. Or you mean  90% of test predictions occur on X-Y locations seen in training **and generated waypoints**?",
    "1313913": "About 85% of our optimisation outcomes are waypoints from the train grid. We make the leap here that the optimization is unbiased. This is likely since we observed no bias in validation.",
    "1313983": "Thanks Tom for sharing the solution. I strongly believe this is one of the greatest solution I saw in Kaggle. Your overall strategy is much better but similar to our strategy, because the heart of our solution is also the discrete optimization and the penalties you use are somewhat similar to the ones we use. \n\nThe reason we lost seems (This is my poem, maybe I should write it in my solution😂, which I started writing just now.):  \n1) The optimization technique (that is used in heuristic contests) is much better than ours. You did beam search, which is much better than the simple greedy algorithm used in our postprocessing. I tried optimization chunk by chunk. \n2) Your grid generation strategy is great. I think probably we would have tried a similar strategy if you didn't hide LB score, but anyway we would have lost because of 1) and 3). \n3) Our absolute position prediction is much better than yours, but MAE of my delta prediction model is about 1.05, which is much worse than yours and it should be more important. \n\nBTW, I have some questions about your waypoint generation techniques. Did you do per-site-floor parameter tuning by hand in waypoint generation? what is the number of the parameter? I thought about the similar techniques, but I thought the number of the parameter will become too big and seems I have to tune the parameters by site-floor (and I was a bit afraid it may be regarded as hand label 😭), so I gave up it.\n\nAnother question is about your teammates. Which parts did your teammates (dott, areeh) focused on? I think your solution is too great and cannot be performed by solo, even if you can work 20 hours per day.",
    "1314007": "Your team's solutions are really interesting, thanks for sharing!\n\nI worked on the grid generation so I can answer your questions there. There are no per-site-floor parameters that are tuned by hand (this to me would be against the spirit of the competition and similar to labeling by hand). The parameters that are tuned interact with the per-site-floor statistics. For instance the parameters that change the generation the most are multipliers on the distance statistics. There are a few parameters that are really just encoding global statistics, like what is the max distance to a wall for a \"point close to the wall\" that we will ever consider. There are a lot of minor parameters that are quite easy to pick, like only merge points if they are at a ratio < 1 of the intended generation distance. There are a few important parameters and many (~20+) minor ones.\n\nI wanted to design a more beautiful solution, but most of the grid generation work was a big effort the last week so I had to build more and design less to get a good result in time.",
    "1314019": "areeh, congrats 1st place and thanks for your comment :) \n```\nThere are no per-site-floor parameters that are tuned by hand.\nThe parameters that are tuned interact with the per-site-floor statistics.\n```\nNice, it sounds very fair and really reasonable :) It's just unbelievable for me you did such a great job in only a week!",
    "1314027": "Thank you. My teammates were doing excellent work so it was important to me to work hard",
    "1314104": "Congrats to the entire \"Track me if you can\" @tvdwiele @dott1718 and @areehdot  on an amazing finish. Great solution write-up- There is a lot of unpack here and I'll need to reread a few times. I had expected the top solutions would approach this as some sort of discrete optimization problem- and your team did an amazing job creating an elegant solution. I'm glad my snap-to-grid notebook was insightful, but if I'm honest I'm sure your team would've found it quickly on your own. Bravo!",
    "1314283": "Kudos to the team. Thank you for sharing the solution. Keep the hard-working going!!! 👍",
    "1314647": "Thanks a lot @mamasinkgs and congrats on the 2nd place and an amazing work you and your team have done! It is particularly interesting to see that we can benchmark the accuracy of our individual models only now, and especially to see there is a large gap for sensor-based model performance. We are committed to share the entire pipeline within a week or two, so you can rerun it and check on your validation data. With regards to absolute position predictions, I was sure we can do better, but all the evidence pointed at them not being that important for the overall picture, so it was our last priority to improve those.",
    "1314669": "Generally speaking, there was no clear separation between who worked on what parts of the project. We very much functioned as a well-oiled team. I have had the pleasure to work with both Dmitry (in Kaggle) and Are (at our Norwegian startup) and we immediately hit the right collaboration dynamic. It can not be stressed enough that everyone worked very hard to be able to do as well as we did. We all worked on data exploration, brainstorming, following up on public discussions and quality checks of our predictions. The list below gives a rough approximation of the major contributors to the components discussed in our writeup:\n\n- **WiFi models**: Dmitry and Are tried NN modeling. Dmitry built the final LightGBM WiFi model and Tom worked on the kNN WiFi model.\n- **Sensor models**: Tom built some low-performant initial sensor models and Dmitry took charge of sensor modeling after that and made massive improvement over the initial models.\n- **Data leaks**: Mostly a collaboration between Dmitry and Tom\n- **Leaderboard probing**: Tom\n- **Waypoint generation**: Are took charge of this component and everything else related to the floor map. This was definitely one of the hardest and most important tasks in this list.\n- **Optimization and ensembling**: Tom",
    "1315143": "Contrats to you guys! This is one of the best solutions i saw here, i can see that a lot of analysis was required to reach up to that score. I'm looking forward to study the code once you guys release it.\n\nThank you!",
    "1315161": "Thank you for sharing the solution!\n\n> The model uses 9 inputs (acce, gyro and ahrs of all 3 axes) and has a relatively simple structure: Conv1d + GRU + Conv1d + head. \n\nCan Cond1d be thought of as a dense layer in this example? Am I correct that for a 10 seconds walk the GRU would be unrolled 10 * 50 (number of IMU recordings per sec) times?",
    "1315263": "A stack of 1D convolution can be thought of as a MLP which is applied to the overlapping time windows and where all weights are shared along the time dimension. [Reference](https://www.quora.com/What-is-the-difference-between-a-convolutional-neural-network-and-a-multilayer-perceptron)\n\nWe would indeed unroll the GRU 500 steps in time in your example.",
    "1316884": "Congratulations to the team and thanks for sharing. Really interesting work!!",
    "1318247": "Thank you for sharing this!",
    "1322952": "Update: We have made all our code available in a [public GitHub repo](https://github.com/ttvand/Indoor-Location-Navigation-Public). Enjoy!",
    "1715034": "Hi there, congratulations. I am a rookie student and happen to see your work. I am interested in your sensor models. I am confused that whether the three sensor models u mentioned in the passage are corresponding to the file named 'sensor_model_movement 1 and 2' and 'sensor_model_dist' in github web page. If yes, what are the usages of files in unused folder for there are some sensor models too which are a little different. Honestly asking for teaching and being appreciated",
    "2135905": "Thank you for sharing! 👍"
  },
  "source": "meta"
}