{
  "id": 240773,
  "title": "15th place solution - Dive into compute_f without Grid Generation + Simple 100% KNN Floor",
  "url": "/competitions/indoor-location-navigation/writeups/algorithm-is-all-you-need-15th-place-solution-dive",
  "author_name": "",
  "post_date": "2022-06-16T15:12:22.440Z",
  "votes": 24,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Thank the host for such an interesting competition. And also thank my great teammates for working hard throughout the competition ( <a href=\"https://www.kaggle.com/ryotayoshinobu\" target=\"_blank\">@ryotayoshinobu</a>, <a href=\"https://www.kaggle.com/tubotubo\" target=\"_blank\">@tubotubo</a>, <a href=\"https://www.kaggle.com/columbia2131\" target=\"_blank\">@columbia2131</a> ).<br>\nThe key factor of this competition is post processing based on algorithms rather than modeling. Also, all of us like competitive programming, so there is no other team’s name but “Algorithm is All You Need” :)</p>\n<p>Here is our solution.</p>\n<p><strong>Quick summary</strong></p>\n<p><img src=\"https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture1.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&amp;X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&amp;X-Goog-Date=20220616T150747Z&amp;X-Goog-Expires=345599&amp;X-Goog-SignedHeaders=host&amp;X-Goog-Signature=7de5d6ffefbd8e3fa8770da60dbacdec1604c7c080f3bf63a729f7466c446bbcc72b2ea6425bd349b5732571440fe2a7856a80873dfdaadbec80d9f1104409bd2940ec32abeae4d516a6f84c5dfa2c25308abdd160a1c199f3c2f693ac7dd3217adf9b4fdb12fdaa8d2369d081338f9cd12552af7a0e0f1f28fc2bdef89b4e43b6163cfb77d2aec9125af12c27874e079d5b99dd1625b6455118cd93c3a0e6c071d0cc2d8f7c7a33c589b0183226cb191a773eb547a23e6f906e56432765827fd0f5156eab0daaae6c528329febf77d2797ad2b128ccbe18ad9df23406520fe921f79ccc2e85d6dd3492d8637b73d74086c4134d4c77608a1a7d61878a38c4bb\" alt=\"\"></p>\n<p>This competition required a good amount of pre and post processing. Our final prediction was a blend of 3rd stage models trained with iterative pseudo labeling. Repeating the post processing was also important for our solution.</p>\n<p><strong>Simple 100% Accurate Floor model</strong></p>\n<ul>\n<li>Consider RSSI as the number of occurrences of a BSSID.</li>\n<li>TF-IDF vectorization.</li>\n<li>KNN for each site.</li>\n</ul>\n<p>Our KNN model performed 100% accuracy for both public and private dataset.</p>\n<p><strong>Waypoint Model</strong></p>\n<p><strong>Preprocess</strong></p>\n<ul>\n<li>Linear interpolation of waypoints based on WiFi timestamp.</li>\n<li>Fix malformed txt data <a href=\"https://www.kaggle.com/higepon\" target=\"_blank\">@higepon</a> ‘s notebook <br>\n<a href=\"https://www.kaggle.com/higepon/how-to-fix-malformed-train-test-data\" target=\"_blank\">https://www.kaggle.com/higepon/how-to-fix-malformed-train-test-data</a></li>\n<li>Remove WiFi information which exists only in training datasets.</li>\n<li>Applying Kalman Filter to sensor data for getting a precise result of compute_f.</li>\n</ul>\n<p><strong>Cross Validation</strong></p>\n<ul>\n<li>GroupKFold of path</li>\n<li>5 CV but n_splits=15 to get many combinations of path groups to enhance random seed averaging.</li>\n</ul>\n<p><strong>1st stage training</strong></p>\n<ul>\n<li>LSTM based on <a href=\"https://www.kaggle.com/Kouki\" target=\"_blank\">@Kouki</a> ‘s notebook <br>\n<a href=\"https://www.kaggle.com/kokitanisaka/lstm-by-keras-with-unified-wi-fi-feats\" target=\"_blank\">https://www.kaggle.com/kokitanisaka/lstm-by-keras-with-unified-wi-fi-feats</a>.<br>\nHyper parameters were tuned with Optuna.</li>\n<li>Transformer-like Convolutional Encoder.<br>\nEmbedding BSSID and RSSI, then Positional Encoding are added.<br>\nNo Feed Forward Network.<br>\nBatch Normalization instead of Layer Normalization.<br>\nMultiplying attention instead of Adding.<br>\nQuery of MultiHeadAttention is BSSID embeddings, Key and Value are RSSI embeddings.<br>\nEncoding process is like [Embedding -&gt; Conv1d -&gt; Multiply Attention -&gt; Conv1d -&gt; Multiply Attention -&gt; … -&gt; Dense]</li>\n<li>10 random seeds to learn multiple combinations of path groups.</li>\n</ul>\n<p><strong>2nd stage training</strong></p>\n<ul>\n<li>Convolutional stacking to learn correlation between different models or different random seeds, which is equal to learning multiple path combinations.</li>\n<li>LightGBM to learn time series and relative position information.<br>\nAdditional input features are:</li>\n</ul>\n<ol>\n<li>n predictions before and after</li>\n<li>Difference between n predictions</li>\n<li>Rate of change from n predictions</li>\n<li>Difference between the previous and next n predictions</li>\n<li>Moving average</li>\n<li>Moving variance</li>\n<li>Relative position (compute_f, compute_rel_position)</li>\n<li>Cumulative sum of relative positions</li>\n<li>Difference between predicted values and cumulative sum of relative positions</li>\n<li>Aggregate features for each path (mean,max,min,median,std,sum)</li>\n</ol>\n<p><strong>3rd stage training</strong></p>\n<ul>\n<li>Hill Climbing to get weights which minimize loss.</li>\n<li>Ridge regression to suppress overfitting.</li>\n</ul>\n<p><strong>Post process</strong><br>\nPost processing is very important to push up scores significantly, and also a fun part for competitive programmers. Here we describe detailed steps. Summary is described below figure.</p>\n<p><img src=\"https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture2.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&amp;X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&amp;X-Goog-Date=20220616T150904Z&amp;X-Goog-Expires=345599&amp;X-Goog-SignedHeaders=host&amp;X-Goog-Signature=937a2b44541155f08662c1318ddbc7617da25e94aaf7e15e38c04ebadbd033350968cc93dbf5fc6ab8443f17d3d9d343ed5363cc3099a77fff710b2ea5d3133b82fc5af18813b6fd4d90cb2940b3036696ad144c3697554209973e0b1b355f164623746548bda9d05330b37b5aed9fdd2ee6a6d89adb9766b7818ea89cd2afddbaad16538844dcf3cc62db42b062c18f6666c395aa55ff71171915760401f39484c470946f39560a56bd67f9a34fe21f6ae8e6e472158e15363516fbd74380b72d324df2d44f54d61798d120e4eefeb444fb8845208de108fdc35f0babfce49bd3150c5f49d38b046350534196041275e4d8b9c97a5d3f50f4c6d61c8f4caf48\" alt=\"\"></p>\n<p><strong>Snap to grid</strong><br>\npublished by <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a><br>\n<a href=\"https://www.kaggle.com/robikscube/indoor-navigation-snap-to-grid-post-processing\" target=\"_blank\">https://www.kaggle.com/robikscube/indoor-navigation-snap-to-grid-post-processing</a><br>\nThe idea is to replace the predicted coordinates with the nearest waypoint of the training data. There are two positive effects:<br>\n  1) pushing the predicted waypoint into the hallway when it is inside an obstacle.<br>\n  2) enables us to predict the same point as the training data.<br>\nAs you all know, there are a lot of overlapping waypoints between the training and the test data. So, selecting the coordinate from train waypoints is very effective. Since our team applied iterative post processing, we set threshold 6 in the first half post-processing parts to increase probability of snapping, and threshold 1.5 in the latter half to avoid strange snapping of already precise waypoints.</p>\n<p><strong>Parallel movement by Snap to grid</strong><br>\nAs mentioned above, Snap to grid snaps predictions to the nearest point. However, the shape of the snapped path will not keep its original shape, especially most of the path is predicted to be inside a wall because snap direction is unpredictable. Therefore, before doing the snap to grid, we consider parallel moving of path to direction which is likely to be snapped while maintaining original shape. We calculate the difference between the original coordinates and the coordinates after Snap to Grid, and regard the average values of difference dx and dy as the direction in which Snap is likely to occur. By adding dx and dy to the original coordinates, we were able to move the path parallelly to the direction that is likely to be snapped while maintaining the original shape of the path.</p>\n<p><strong>Cost Minimization</strong><br>\npublished by <a href=\"https://www.kaggle.com/saitodevel01\" target=\"_blank\">@saitodevel01</a><br>\n<a href=\"https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization\" target=\"_blank\">https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization</a></p>\n<p>The idea is to minimize the difference between the distance to the predicted point and the distance calculated by the sensor data.</p>\n<p>⊿X^ is calculated by compute_rel_positions in compute_f. Though the host's prediction of relative coordinates with compute_rel_positions function is reasonably accurate, its accuracy depends on the quality of raw sensor data. If you look carefully at the relative coordinates, you will notice that the path of relative coordinates has been scaled up and that there is a bias in the direction of rotation compared to the original path. Let’s check the precision of this function.</p>\n<p><img src=\"https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture3.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&amp;X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&amp;X-Goog-Date=20220616T150947Z&amp;X-Goog-Expires=345599&amp;X-Goog-SignedHeaders=host&amp;X-Goog-Signature=07cf865742e7f372aab6e278bc1b695315606c4104818bd9c1b04052c23d464e341c17ee6a361e5b39c893c3a67a438b25c178a6d674de09bf5d0d439df37076006e13f53db2ded49a9e2ed5da4d9f137d75bb6d18feee313799f4909651e27ccb850d8e6fe1a3cd850247a7cf71409c1b7692326aa7da8b3def2c683bae7cc63702fc2718ae8cbfc26faa80748c8be2df222df13da5736ee061996cd459979c638f5e4d907587124d00a7719f2ca18ca3e47518f408c8aec2434fc5135c6deb33558a5a6337b738ec1918076e54d96786a4e9b6acc09fa2905753d4d285404f6b42930196db37583d2dc63884a57ab73879718efa85b43501731c627ba2deea\" alt=\"\"></p>\n<p>Blue line indicates training path, and red line is relative positions. It is clear that relative position is scaled up and rotated compared to red line, so we need to calibrate compute_f function to get a good relative position. Compute_f function uses sensor data, so we try to clean up raw sensor data by Kalman Filter. Using Kalman filtered sensor data produces nice relative positions. We can get similar scale relative positions similar to the original path, which leads to a good score by cost minimization.</p>\n<p><img src=\"https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture4.png?X-Goog-Algorithm=GOOG4-RSA-SHA256&amp;X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&amp;X-Goog-Date=20220616T151012Z&amp;X-Goog-Expires=345599&amp;X-Goog-SignedHeaders=host&amp;X-Goog-Signature=198c723ffdd4fdb1de79a1deacda24fa8606b2d5ba0eb809ba0a90e4b22fdd237c1ae08f1b3bb261f2713e472f88c21cdc8a6250a53392398ba9b3702b50708dd602e4a2ad1328d5dac30e6e4d3eca7fe90a5ce7a269dece0ed30fcc4e3d94c78831586c666f158a8d211ea60d15f0c2c7016e93acd0f7b8ac79d30fb91727bed9f4383849659a3a84110274d0e7b230ff52c63d5eb9d0794af0ea31ba8ead3e7106d92142c9a46ee2682cee211dd17afb948f8432a143995067e6c06d4f06c0471ef8b529c8f84ccac431deb10d62cd58ba0bdcbe09ef66ba4b4d5177dd3b94f95a9ff792c1b9d39e967eba1a1e5efcb0078fbe34050325ab7c908e1f04d40b\" alt=\"\"></p>\n<p>Next, we tackle with the rotation problems. Here is an example.</p>\n<p><img src=\"https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture5.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&amp;X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&amp;X-Goog-Date=20220616T151036Z&amp;X-Goog-Expires=345599&amp;X-Goog-SignedHeaders=host&amp;X-Goog-Signature=32e5d5e7c36cc9354a4c68341ac085a58c5b85953efe3e4980bcca888b514d74dbc114d916b4080ce7e258f55f8d9a5e57d495bed186d4cf0c7c0540528b245d4fa9caf36bf7347e5cdf18f7d05afb436a3aaffddef9c4f1edda32d0757096fdba1113e985423fa93dc314cfa078c600691740dbefcb38b3f9e0469b7ed87e41a04263dde99c7c098fb4b464001796ecf1587edaede52fa2a7019f00b5e74ff2cdaf14a7100d7b9d8a49ab38120f87eb83d2389d62aaba0856738e5b170f45e34df8d584b2263179b1129bb477db336be58fdc7ef1954c84e56f61663dde341d1bcc1008ab96b05f2bb522967f2e4963696d1934ec17160aa51b5892a78df5dd\" alt=\"\"></p>\n<p>It is obvious that the red lines are rotated some degree. We try to solve these rotation problems with minimization of Euclid distance. </p>\n<p>Let the predicted coordinates in a certain path after post-processing be A1, A2, … An in order from the start point, and the predicted coordinates by compute_rel_positions be B1, B2 … Bn.If there is no rotation, the movement amount delta_a of A1 → A2 and the movement amount delta_b of B1 → B2 should be roughly the same. However, as you can see from the visualization, compute_rel_positions will rotate by some degrees, so delta_a and delta_b do not match. Therefore, we tried brute force search for the rotation angle θ of delta_b which minimizes the difference between delta_a and delta_b.<br>\nThe loss function is the sum of the distances between delta_a and delta_b.</p>\n<p>loss (θ) = || (A2-A1) – (B2-B1) R (θ) || ^ 2 + || (A3-A2)-(B3-B2) R (θ) || ^ 2 +… + || (An-An-1) – (Bn-Bn-1) R (θ) || ^ 2</p>\n<p>Where R (θ) is a rotation matrix that rotates the coordinates by θ degrees. For example, (B3 – B2) R (θ) represents B2 → B3 rotated by θ. Also, || X – Y || ^ 2 represents the square of the difference between the movements X and Y.In other words, the loss function represents the difference between the \"current movement amount\" and “the movement amount calculated by compute_rel_positions rotated by θ”.The smaller the value of the loss function, the closer the angle of B is to A.</p>\n<p><img src=\"https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture6.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&amp;X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&amp;X-Goog-Date=20220616T151051Z&amp;X-Goog-Expires=345599&amp;X-Goog-SignedHeaders=host&amp;X-Goog-Signature=7fdaab9ded4bacb795239e4200aa3bbeff897865c5177983910adadc6d812262c9006766583e93310dec882e950b3199632f76b9fba5d63697bb806f39716ffe17f23c78384d4adca42b672263802c1835b8bcfd51fa95931b9765dc13fe8a2477f3ce957016696f3e01353f1612b41ec544fe220117fac3d302ee656737966f4460ccf423eb6b0f0b0e4097596ddf75711554448bca8e8b3fb1c544d011b15253bc98c97e04fa5388f39f1a3b7469007cd73951ff74acd0b8faffc723e2345c76abcd168b515699380cc3bb1dcfc25a03bf64373a926d1eaf7ee9f498df786848f7461dff5b7d0c079a4d36f871e5f36827192a0796705b4b4f57453658c412\" alt=\"\"></p>\n<p>We modified the submission path by compute_rel_positions in the post process pipeline, and modified the rotation angle of compute_rel_positions by submission at the end of the pipeline. As we repeated the post process pipeline, both submission path and compute_rel_positions enhanced each other.</p>\n<p>The difference between original and modified compute_f are very clear.</p>\n<p><img src=\"https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture7.png?X-Goog-Algorithm=GOOG4-RSA-SHA256&amp;X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&amp;X-Goog-Date=20220616T151109Z&amp;X-Goog-Expires=345599&amp;X-Goog-SignedHeaders=host&amp;X-Goog-Signature=3df41bfcad96849247ca344feda8468daa5b7569fb560d1df7fea21504685437ac7dcd56a8393b251b61e4e2af32ab8824ab0df56476e90965f8f0286ff0e8fd0f0bb775da3b46fc93a9495030c44a5542679ff0a2cb7e14106b92fdeb36cf0e340c75a3b3960b4cfbf3f68b1ab938945060bf9f86d17a13dac01bdba56e85a6d7dcd7dcc1b86ebaa7bd3e016a0f279d60d3f55189eba80a48e00afd71cb04d07a1284508fce1b92372624c4e9099298f24dfcf431d42da59fb98a503b726e75c952281089d4715e3289003cc3209400cd95d5406556cade8d14be1067366f8427b53ac3e7e4054432bf67b7c6f084de8047513bfc58ed94804afb0b5caf3203\" alt=\"\"></p>\n<p><img src=\"https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture8.png?X-Goog-Algorithm=GOOG4-RSA-SHA256&amp;X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&amp;X-Goog-Date=20220616T151123Z&amp;X-Goog-Expires=345599&amp;X-Goog-SignedHeaders=host&amp;X-Goog-Signature=b073cb1f4daffe9761010eefa4cf2387650b188f410fec31029d00f163b6575e7e0f7bc3e5cf2987c234e355a14bebe892192afeac637262d15ef934ce594085df4960184a9ad7e4c988cc119708c5349192b808dc1d4b47d0522bb6b50a2647626ebe66e17db70b71f3222ea2774f4a24a29ddb1f97710d93f0ff99ba57ddb4e2201a71ef755aad92720bc8cf7e2d42b62e25abf985e2b5498e2193c1b9fcc00b4cdd8f4c563d0687f7694cb0e49ad411a7e915880aa435ef4dd03e38040782aef7f8ed90777e16495ae8407011483427fe00325de9f5ef146be85202f417a1c6ab5241afa9ab3555e842b31694f2b299a210f2df5f172ff847a903ff76ac82\" alt=\"\"></p>\n<p>In addition, cost minimization has two important parameters, α and β. The higher the α, the more emphasize the current prediction, and the higher the β, the more emphasize the sensor data. Since we use Cost Minimization many times in our pipeline, β was attenuated with each successive use so that Cost Minimization would not be too influenced by the sensor data.</p>\n<p><strong>Push to Hallway</strong><br>\nIf predictions are located in an obstacle, we need to push them to the nearest waypoint. But when we just increase the threshold of Snap to Grid, waypoints are changed to strange waypoints since there are many hallways which don’t have waypoints. So, we tried to push the waypoints to the nearest hallway even if there were no waypoints.<br>\nEach pixel of the floor image was used to determine whether predictions are located in the hallway or an obstacle, and moved waypoints located in the obstacle to the nearest coordinates in the hallway. This has the great advantage of being able to modify the waypoints in obstacles which is far from training waypoints.</p>\n<p><img src=\"https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture9.png?X-Goog-Algorithm=GOOG4-RSA-SHA256&amp;X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&amp;X-Goog-Date=20220616T151136Z&amp;X-Goog-Expires=345599&amp;X-Goog-SignedHeaders=host&amp;X-Goog-Signature=23a9f942b3a842db1df2d1ab7a051ffca9450987bcc796b7c492e8749d40b8845c41f32842723abe09e4ca53a2fa53a5894a235f1108fb4f9a175653249bbda8ef80d92406c954508358c822c29dbfd77575541191e5082c2ad71a1fd233f98a6ad11f396e1e38f60519adc63520727f9e6c85e7af77707f72273e3fc0f11eebf0e82b7d31bf3c40b38ce95ccf4fc21813e23445ad1388fcd6b7fd5a7b4cf98a6d87a5e8b125e11b6fb2b216df689a1fed4e0937f22841ba5e3aa3c7ebab450de19305b2bd3ebcf2ddd12ba451bcf74eb89b50df135176f115d7133d915af0f79472a286b609ed3d542616bc5f471da8abfeac66b146f1f29345ae033550c1b1\" alt=\"\"></p>\n<p><strong>Apply Start and End Points leakage</strong><br>\npublished by <a href=\"https://www.kaggle.com/tomooinubushi\" target=\"_blank\">@tomooinubushi</a><br>\n<a href=\"https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage\" target=\"_blank\">https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage</a></p>\n<p>This was beyond our analytical capabilities, so we only corrected the start and end waypoints as same as the public notebook.</p>\n<p><strong>What didn’t work</strong><br>\nTime series RNN.<br>\nInput all WiFi and sensor data to NN.<br>\neach model for each site.<br>\nPre training of all data and Transfer Learning.<br>\nDijkstra, Warshall-Froyd, Band First Search as shortest path solving <br>\nMap Matching<br>\nCurve to Curve<br>\nGrid Generation</p>\n<p>Thank you for reading our solution. <br>\nAny opinions are welcome!</p>",
  "messages": [
    {
      "id": "1317353",
      "postDate": "05/21/2021 10:57:07",
      "content": "<p>Thank the host for such an interesting competition. And also thank my great teammates for working hard throughout the competition ( <a href=\"https://www.kaggle.com/ryotayoshinobu\" target=\"_blank\">@ryotayoshinobu</a>, <a href=\"https://www.kaggle.com/tubotubo\" target=\"_blank\">@tubotubo</a>, <a href=\"https://www.kaggle.com/columbia2131\" target=\"_blank\">@columbia2131</a> ).<br>\nThe key factor of this competition is post processing based on algorithms rather than modeling. Also, all of us like competitive programming, so there is no other team’s name but “Algorithm is All You Need” :)</p>\n<p>Here is our solution.</p>\n<p><strong>Quick summary</strong></p>\n<p><img src=\"https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture1.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&amp;X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&amp;X-Goog-Date=20220616T150747Z&amp;X-Goog-Expires=345599&amp;X-Goog-SignedHeaders=host&amp;X-Goog-Signature=7de5d6ffefbd8e3fa8770da60dbacdec1604c7c080f3bf63a729f7466c446bbcc72b2ea6425bd349b5732571440fe2a7856a80873dfdaadbec80d9f1104409bd2940ec32abeae4d516a6f84c5dfa2c25308abdd160a1c199f3c2f693ac7dd3217adf9b4fdb12fdaa8d2369d081338f9cd12552af7a0e0f1f28fc2bdef89b4e43b6163cfb77d2aec9125af12c27874e079d5b99dd1625b6455118cd93c3a0e6c071d0cc2d8f7c7a33c589b0183226cb191a773eb547a23e6f906e56432765827fd0f5156eab0daaae6c528329febf77d2797ad2b128ccbe18ad9df23406520fe921f79ccc2e85d6dd3492d8637b73d74086c4134d4c77608a1a7d61878a38c4bb\" alt=\"\"></p>\n<p>This competition required a good amount of pre and post processing. Our final prediction was a blend of 3rd stage models trained with iterative pseudo labeling. Repeating the post processing was also important for our solution.</p>\n<p><strong>Simple 100% Accurate Floor model</strong></p>\n<ul>\n<li>Consider RSSI as the number of occurrences of a BSSID.</li>\n<li>TF-IDF vectorization.</li>\n<li>KNN for each site.</li>\n</ul>\n<p>Our KNN model performed 100% accuracy for both public and private dataset.</p>\n<p><strong>Waypoint Model</strong></p>\n<p><strong>Preprocess</strong></p>\n<ul>\n<li>Linear interpolation of waypoints based on WiFi timestamp.</li>\n<li>Fix malformed txt data <a href=\"https://www.kaggle.com/higepon\" target=\"_blank\">@higepon</a> ‘s notebook <br>\n<a href=\"https://www.kaggle.com/higepon/how-to-fix-malformed-train-test-data\" target=\"_blank\">https://www.kaggle.com/higepon/how-to-fix-malformed-train-test-data</a></li>\n<li>Remove WiFi information which exists only in training datasets.</li>\n<li>Applying Kalman Filter to sensor data for getting a precise result of compute_f.</li>\n</ul>\n<p><strong>Cross Validation</strong></p>\n<ul>\n<li>GroupKFold of path</li>\n<li>5 CV but n_splits=15 to get many combinations of path groups to enhance random seed averaging.</li>\n</ul>\n<p><strong>1st stage training</strong></p>\n<ul>\n<li>LSTM based on <a href=\"https://www.kaggle.com/Kouki\" target=\"_blank\">@Kouki</a> ‘s notebook <br>\n<a href=\"https://www.kaggle.com/kokitanisaka/lstm-by-keras-with-unified-wi-fi-feats\" target=\"_blank\">https://www.kaggle.com/kokitanisaka/lstm-by-keras-with-unified-wi-fi-feats</a>.<br>\nHyper parameters were tuned with Optuna.</li>\n<li>Transformer-like Convolutional Encoder.<br>\nEmbedding BSSID and RSSI, then Positional Encoding are added.<br>\nNo Feed Forward Network.<br>\nBatch Normalization instead of Layer Normalization.<br>\nMultiplying attention instead of Adding.<br>\nQuery of MultiHeadAttention is BSSID embeddings, Key and Value are RSSI embeddings.<br>\nEncoding process is like [Embedding -&gt; Conv1d -&gt; Multiply Attention -&gt; Conv1d -&gt; Multiply Attention -&gt; … -&gt; Dense]</li>\n<li>10 random seeds to learn multiple combinations of path groups.</li>\n</ul>\n<p><strong>2nd stage training</strong></p>\n<ul>\n<li>Convolutional stacking to learn correlation between different models or different random seeds, which is equal to learning multiple path combinations.</li>\n<li>LightGBM to learn time series and relative position information.<br>\nAdditional input features are:</li>\n</ul>\n<ol>\n<li>n predictions before and after</li>\n<li>Difference between n predictions</li>\n<li>Rate of change from n predictions</li>\n<li>Difference between the previous and next n predictions</li>\n<li>Moving average</li>\n<li>Moving variance</li>\n<li>Relative position (compute_f, compute_rel_position)</li>\n<li>Cumulative sum of relative positions</li>\n<li>Difference between predicted values and cumulative sum of relative positions</li>\n<li>Aggregate features for each path (mean,max,min,median,std,sum)</li>\n</ol>\n<p><strong>3rd stage training</strong></p>\n<ul>\n<li>Hill Climbing to get weights which minimize loss.</li>\n<li>Ridge regression to suppress overfitting.</li>\n</ul>\n<p><strong>Post process</strong><br>\nPost processing is very important to push up scores significantly, and also a fun part for competitive programmers. Here we describe detailed steps. Summary is described below figure.</p>\n<p><img src=\"https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture2.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&amp;X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&amp;X-Goog-Date=20220616T150904Z&amp;X-Goog-Expires=345599&amp;X-Goog-SignedHeaders=host&amp;X-Goog-Signature=937a2b44541155f08662c1318ddbc7617da25e94aaf7e15e38c04ebadbd033350968cc93dbf5fc6ab8443f17d3d9d343ed5363cc3099a77fff710b2ea5d3133b82fc5af18813b6fd4d90cb2940b3036696ad144c3697554209973e0b1b355f164623746548bda9d05330b37b5aed9fdd2ee6a6d89adb9766b7818ea89cd2afddbaad16538844dcf3cc62db42b062c18f6666c395aa55ff71171915760401f39484c470946f39560a56bd67f9a34fe21f6ae8e6e472158e15363516fbd74380b72d324df2d44f54d61798d120e4eefeb444fb8845208de108fdc35f0babfce49bd3150c5f49d38b046350534196041275e4d8b9c97a5d3f50f4c6d61c8f4caf48\" alt=\"\"></p>\n<p><strong>Snap to grid</strong><br>\npublished by <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a><br>\n<a href=\"https://www.kaggle.com/robikscube/indoor-navigation-snap-to-grid-post-processing\" target=\"_blank\">https://www.kaggle.com/robikscube/indoor-navigation-snap-to-grid-post-processing</a><br>\nThe idea is to replace the predicted coordinates with the nearest waypoint of the training data. There are two positive effects:<br>\n  1) pushing the predicted waypoint into the hallway when it is inside an obstacle.<br>\n  2) enables us to predict the same point as the training data.<br>\nAs you all know, there are a lot of overlapping waypoints between the training and the test data. So, selecting the coordinate from train waypoints is very effective. Since our team applied iterative post processing, we set threshold 6 in the first half post-processing parts to increase probability of snapping, and threshold 1.5 in the latter half to avoid strange snapping of already precise waypoints.</p>\n<p><strong>Parallel movement by Snap to grid</strong><br>\nAs mentioned above, Snap to grid snaps predictions to the nearest point. However, the shape of the snapped path will not keep its original shape, especially most of the path is predicted to be inside a wall because snap direction is unpredictable. Therefore, before doing the snap to grid, we consider parallel moving of path to direction which is likely to be snapped while maintaining original shape. We calculate the difference between the original coordinates and the coordinates after Snap to Grid, and regard the average values of difference dx and dy as the direction in which Snap is likely to occur. By adding dx and dy to the original coordinates, we were able to move the path parallelly to the direction that is likely to be snapped while maintaining the original shape of the path.</p>\n<p><strong>Cost Minimization</strong><br>\npublished by <a href=\"https://www.kaggle.com/saitodevel01\" target=\"_blank\">@saitodevel01</a><br>\n<a href=\"https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization\" target=\"_blank\">https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization</a></p>\n<p>The idea is to minimize the difference between the distance to the predicted point and the distance calculated by the sensor data.</p>\n<p>⊿X^ is calculated by compute_rel_positions in compute_f. Though the host's prediction of relative coordinates with compute_rel_positions function is reasonably accurate, its accuracy depends on the quality of raw sensor data. If you look carefully at the relative coordinates, you will notice that the path of relative coordinates has been scaled up and that there is a bias in the direction of rotation compared to the original path. Let’s check the precision of this function.</p>\n<p><img src=\"https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture3.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&amp;X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&amp;X-Goog-Date=20220616T150947Z&amp;X-Goog-Expires=345599&amp;X-Goog-SignedHeaders=host&amp;X-Goog-Signature=07cf865742e7f372aab6e278bc1b695315606c4104818bd9c1b04052c23d464e341c17ee6a361e5b39c893c3a67a438b25c178a6d674de09bf5d0d439df37076006e13f53db2ded49a9e2ed5da4d9f137d75bb6d18feee313799f4909651e27ccb850d8e6fe1a3cd850247a7cf71409c1b7692326aa7da8b3def2c683bae7cc63702fc2718ae8cbfc26faa80748c8be2df222df13da5736ee061996cd459979c638f5e4d907587124d00a7719f2ca18ca3e47518f408c8aec2434fc5135c6deb33558a5a6337b738ec1918076e54d96786a4e9b6acc09fa2905753d4d285404f6b42930196db37583d2dc63884a57ab73879718efa85b43501731c627ba2deea\" alt=\"\"></p>\n<p>Blue line indicates training path, and red line is relative positions. It is clear that relative position is scaled up and rotated compared to red line, so we need to calibrate compute_f function to get a good relative position. Compute_f function uses sensor data, so we try to clean up raw sensor data by Kalman Filter. Using Kalman filtered sensor data produces nice relative positions. We can get similar scale relative positions similar to the original path, which leads to a good score by cost minimization.</p>\n<p><img src=\"https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture4.png?X-Goog-Algorithm=GOOG4-RSA-SHA256&amp;X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&amp;X-Goog-Date=20220616T151012Z&amp;X-Goog-Expires=345599&amp;X-Goog-SignedHeaders=host&amp;X-Goog-Signature=198c723ffdd4fdb1de79a1deacda24fa8606b2d5ba0eb809ba0a90e4b22fdd237c1ae08f1b3bb261f2713e472f88c21cdc8a6250a53392398ba9b3702b50708dd602e4a2ad1328d5dac30e6e4d3eca7fe90a5ce7a269dece0ed30fcc4e3d94c78831586c666f158a8d211ea60d15f0c2c7016e93acd0f7b8ac79d30fb91727bed9f4383849659a3a84110274d0e7b230ff52c63d5eb9d0794af0ea31ba8ead3e7106d92142c9a46ee2682cee211dd17afb948f8432a143995067e6c06d4f06c0471ef8b529c8f84ccac431deb10d62cd58ba0bdcbe09ef66ba4b4d5177dd3b94f95a9ff792c1b9d39e967eba1a1e5efcb0078fbe34050325ab7c908e1f04d40b\" alt=\"\"></p>\n<p>Next, we tackle with the rotation problems. Here is an example.</p>\n<p><img src=\"https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture5.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&amp;X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&amp;X-Goog-Date=20220616T151036Z&amp;X-Goog-Expires=345599&amp;X-Goog-SignedHeaders=host&amp;X-Goog-Signature=32e5d5e7c36cc9354a4c68341ac085a58c5b85953efe3e4980bcca888b514d74dbc114d916b4080ce7e258f55f8d9a5e57d495bed186d4cf0c7c0540528b245d4fa9caf36bf7347e5cdf18f7d05afb436a3aaffddef9c4f1edda32d0757096fdba1113e985423fa93dc314cfa078c600691740dbefcb38b3f9e0469b7ed87e41a04263dde99c7c098fb4b464001796ecf1587edaede52fa2a7019f00b5e74ff2cdaf14a7100d7b9d8a49ab38120f87eb83d2389d62aaba0856738e5b170f45e34df8d584b2263179b1129bb477db336be58fdc7ef1954c84e56f61663dde341d1bcc1008ab96b05f2bb522967f2e4963696d1934ec17160aa51b5892a78df5dd\" alt=\"\"></p>\n<p>It is obvious that the red lines are rotated some degree. We try to solve these rotation problems with minimization of Euclid distance. </p>\n<p>Let the predicted coordinates in a certain path after post-processing be A1, A2, … An in order from the start point, and the predicted coordinates by compute_rel_positions be B1, B2 … Bn.If there is no rotation, the movement amount delta_a of A1 → A2 and the movement amount delta_b of B1 → B2 should be roughly the same. However, as you can see from the visualization, compute_rel_positions will rotate by some degrees, so delta_a and delta_b do not match. Therefore, we tried brute force search for the rotation angle θ of delta_b which minimizes the difference between delta_a and delta_b.<br>\nThe loss function is the sum of the distances between delta_a and delta_b.</p>\n<p>loss (θ) = || (A2-A1) – (B2-B1) R (θ) || ^ 2 + || (A3-A2)-(B3-B2) R (θ) || ^ 2 +… + || (An-An-1) – (Bn-Bn-1) R (θ) || ^ 2</p>\n<p>Where R (θ) is a rotation matrix that rotates the coordinates by θ degrees. For example, (B3 – B2) R (θ) represents B2 → B3 rotated by θ. Also, || X – Y || ^ 2 represents the square of the difference between the movements X and Y.In other words, the loss function represents the difference between the \"current movement amount\" and “the movement amount calculated by compute_rel_positions rotated by θ”.The smaller the value of the loss function, the closer the angle of B is to A.</p>\n<p><img src=\"https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture6.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&amp;X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&amp;X-Goog-Date=20220616T151051Z&amp;X-Goog-Expires=345599&amp;X-Goog-SignedHeaders=host&amp;X-Goog-Signature=7fdaab9ded4bacb795239e4200aa3bbeff897865c5177983910adadc6d812262c9006766583e93310dec882e950b3199632f76b9fba5d63697bb806f39716ffe17f23c78384d4adca42b672263802c1835b8bcfd51fa95931b9765dc13fe8a2477f3ce957016696f3e01353f1612b41ec544fe220117fac3d302ee656737966f4460ccf423eb6b0f0b0e4097596ddf75711554448bca8e8b3fb1c544d011b15253bc98c97e04fa5388f39f1a3b7469007cd73951ff74acd0b8faffc723e2345c76abcd168b515699380cc3bb1dcfc25a03bf64373a926d1eaf7ee9f498df786848f7461dff5b7d0c079a4d36f871e5f36827192a0796705b4b4f57453658c412\" alt=\"\"></p>\n<p>We modified the submission path by compute_rel_positions in the post process pipeline, and modified the rotation angle of compute_rel_positions by submission at the end of the pipeline. As we repeated the post process pipeline, both submission path and compute_rel_positions enhanced each other.</p>\n<p>The difference between original and modified compute_f are very clear.</p>\n<p><img src=\"https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture7.png?X-Goog-Algorithm=GOOG4-RSA-SHA256&amp;X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&amp;X-Goog-Date=20220616T151109Z&amp;X-Goog-Expires=345599&amp;X-Goog-SignedHeaders=host&amp;X-Goog-Signature=3df41bfcad96849247ca344feda8468daa5b7569fb560d1df7fea21504685437ac7dcd56a8393b251b61e4e2af32ab8824ab0df56476e90965f8f0286ff0e8fd0f0bb775da3b46fc93a9495030c44a5542679ff0a2cb7e14106b92fdeb36cf0e340c75a3b3960b4cfbf3f68b1ab938945060bf9f86d17a13dac01bdba56e85a6d7dcd7dcc1b86ebaa7bd3e016a0f279d60d3f55189eba80a48e00afd71cb04d07a1284508fce1b92372624c4e9099298f24dfcf431d42da59fb98a503b726e75c952281089d4715e3289003cc3209400cd95d5406556cade8d14be1067366f8427b53ac3e7e4054432bf67b7c6f084de8047513bfc58ed94804afb0b5caf3203\" alt=\"\"></p>\n<p><img src=\"https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture8.png?X-Goog-Algorithm=GOOG4-RSA-SHA256&amp;X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&amp;X-Goog-Date=20220616T151123Z&amp;X-Goog-Expires=345599&amp;X-Goog-SignedHeaders=host&amp;X-Goog-Signature=b073cb1f4daffe9761010eefa4cf2387650b188f410fec31029d00f163b6575e7e0f7bc3e5cf2987c234e355a14bebe892192afeac637262d15ef934ce594085df4960184a9ad7e4c988cc119708c5349192b808dc1d4b47d0522bb6b50a2647626ebe66e17db70b71f3222ea2774f4a24a29ddb1f97710d93f0ff99ba57ddb4e2201a71ef755aad92720bc8cf7e2d42b62e25abf985e2b5498e2193c1b9fcc00b4cdd8f4c563d0687f7694cb0e49ad411a7e915880aa435ef4dd03e38040782aef7f8ed90777e16495ae8407011483427fe00325de9f5ef146be85202f417a1c6ab5241afa9ab3555e842b31694f2b299a210f2df5f172ff847a903ff76ac82\" alt=\"\"></p>\n<p>In addition, cost minimization has two important parameters, α and β. The higher the α, the more emphasize the current prediction, and the higher the β, the more emphasize the sensor data. Since we use Cost Minimization many times in our pipeline, β was attenuated with each successive use so that Cost Minimization would not be too influenced by the sensor data.</p>\n<p><strong>Push to Hallway</strong><br>\nIf predictions are located in an obstacle, we need to push them to the nearest waypoint. But when we just increase the threshold of Snap to Grid, waypoints are changed to strange waypoints since there are many hallways which don’t have waypoints. So, we tried to push the waypoints to the nearest hallway even if there were no waypoints.<br>\nEach pixel of the floor image was used to determine whether predictions are located in the hallway or an obstacle, and moved waypoints located in the obstacle to the nearest coordinates in the hallway. This has the great advantage of being able to modify the waypoints in obstacles which is far from training waypoints.</p>\n<p><img src=\"https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture9.png?X-Goog-Algorithm=GOOG4-RSA-SHA256&amp;X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&amp;X-Goog-Date=20220616T151136Z&amp;X-Goog-Expires=345599&amp;X-Goog-SignedHeaders=host&amp;X-Goog-Signature=23a9f942b3a842db1df2d1ab7a051ffca9450987bcc796b7c492e8749d40b8845c41f32842723abe09e4ca53a2fa53a5894a235f1108fb4f9a175653249bbda8ef80d92406c954508358c822c29dbfd77575541191e5082c2ad71a1fd233f98a6ad11f396e1e38f60519adc63520727f9e6c85e7af77707f72273e3fc0f11eebf0e82b7d31bf3c40b38ce95ccf4fc21813e23445ad1388fcd6b7fd5a7b4cf98a6d87a5e8b125e11b6fb2b216df689a1fed4e0937f22841ba5e3aa3c7ebab450de19305b2bd3ebcf2ddd12ba451bcf74eb89b50df135176f115d7133d915af0f79472a286b609ed3d542616bc5f471da8abfeac66b146f1f29345ae033550c1b1\" alt=\"\"></p>\n<p><strong>Apply Start and End Points leakage</strong><br>\npublished by <a href=\"https://www.kaggle.com/tomooinubushi\" target=\"_blank\">@tomooinubushi</a><br>\n<a href=\"https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage\" target=\"_blank\">https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage</a></p>\n<p>This was beyond our analytical capabilities, so we only corrected the start and end waypoints as same as the public notebook.</p>\n<p><strong>What didn’t work</strong><br>\nTime series RNN.<br>\nInput all WiFi and sensor data to NN.<br>\neach model for each site.<br>\nPre training of all data and Transfer Learning.<br>\nDijkstra, Warshall-Froyd, Band First Search as shortest path solving <br>\nMap Matching<br>\nCurve to Curve<br>\nGrid Generation</p>\n<p>Thank you for reading our solution. <br>\nAny opinions are welcome!</p>",
      "rawMarkdown": "Thank the host for such an interesting competition. And also thank my great teammates for working hard throughout the competition ( @ryotayoshinobu, @tubotubo, @columbia2131 ).\nThe key factor of this competition is post processing based on algorithms rather than modeling. Also, all of us like competitive programming, so there is no other team’s name but “Algorithm is All You Need” :)\n \nHere is our solution.\n\n\n**Quick summary**\n\n\n![](https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture1.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20220616T150747Z&X-Goog-Expires=345599&X-Goog-SignedHeaders=host&X-Goog-Signature=7de5d6ffefbd8e3fa8770da60dbacdec1604c7c080f3bf63a729f7466c446bbcc72b2ea6425bd349b5732571440fe2a7856a80873dfdaadbec80d9f1104409bd2940ec32abeae4d516a6f84c5dfa2c25308abdd160a1c199f3c2f693ac7dd3217adf9b4fdb12fdaa8d2369d081338f9cd12552af7a0e0f1f28fc2bdef89b4e43b6163cfb77d2aec9125af12c27874e079d5b99dd1625b6455118cd93c3a0e6c071d0cc2d8f7c7a33c589b0183226cb191a773eb547a23e6f906e56432765827fd0f5156eab0daaae6c528329febf77d2797ad2b128ccbe18ad9df23406520fe921f79ccc2e85d6dd3492d8637b73d74086c4134d4c77608a1a7d61878a38c4bb)\n\nThis competition required a good amount of pre and post processing. Our final prediction was a blend of 3rd stage models trained with iterative pseudo labeling. Repeating the post processing was also important for our solution.\n\n**Simple 100% Accurate Floor model**\n\n- Consider RSSI as the number of occurrences of a BSSID.\n- TF-IDF vectorization.\n- KNN for each site.\n\nOur KNN model performed 100% accuracy for both public and private dataset.\n\n\n**Waypoint Model**\n\n**Preprocess**\n-   \tLinear interpolation of waypoints based on WiFi timestamp.\n-   \tFix malformed txt data @higepon ‘s notebook \nhttps://www.kaggle.com/higepon/how-to-fix-malformed-train-test-data\n-   \tRemove WiFi information which exists only in training datasets.\n-   \tApplying Kalman Filter to sensor data for getting a precise result of compute_f.\n\n\n\n**Cross Validation**\n-   \tGroupKFold of path\n-   \t5 CV but n_splits=15 to get many combinations of path groups to enhance random seed averaging.\n\n\n\n**1st stage training**\n-   \tLSTM based on @Kouki ‘s notebook \nhttps://www.kaggle.com/kokitanisaka/lstm-by-keras-with-unified-wi-fi-feats.\nHyper parameters were tuned with Optuna.\n-   \tTransformer-like Convolutional Encoder.\nEmbedding BSSID and RSSI, then Positional Encoding are added.\nNo Feed Forward Network.\nBatch Normalization instead of Layer Normalization.\nMultiplying attention instead of Adding.\nQuery of MultiHeadAttention is BSSID embeddings, Key and Value are RSSI embeddings.\nEncoding process is like [Embedding -> Conv1d -> Multiply Attention -> Conv1d -> Multiply Attention -> … -> Dense]\n-   \t10 random seeds to learn multiple combinations of path groups.\n\n\n\n**2nd stage training**\n-   \tConvolutional stacking to learn correlation between different models or different random seeds, which is equal to learning multiple path combinations.\n-   \tLightGBM to learn time series and relative position information.\nAdditional input features are:\n1. n predictions before and after\n2. Difference between n predictions\n3. Rate of change from n predictions\n4. Difference between the previous and next n predictions\n5. Moving average\n6. Moving variance\n7. Relative position (compute_f, compute_rel_position)\n8. Cumulative sum of relative positions\n9. Difference between predicted values and cumulative sum of relative positions\n10. Aggregate features for each path (mean,max,min,median,std,sum)\n \n\n\n**3rd stage training**\n-   \tHill Climbing to get weights which minimize loss.\n-   \tRidge regression to suppress overfitting.\n \n \n**Post process**\nPost processing is very important to push up scores significantly, and also a fun part for competitive programmers. Here we describe detailed steps. Summary is described below figure.\n\n\n![](https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture2.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20220616T150904Z&X-Goog-Expires=345599&X-Goog-SignedHeaders=host&X-Goog-Signature=937a2b44541155f08662c1318ddbc7617da25e94aaf7e15e38c04ebadbd033350968cc93dbf5fc6ab8443f17d3d9d343ed5363cc3099a77fff710b2ea5d3133b82fc5af18813b6fd4d90cb2940b3036696ad144c3697554209973e0b1b355f164623746548bda9d05330b37b5aed9fdd2ee6a6d89adb9766b7818ea89cd2afddbaad16538844dcf3cc62db42b062c18f6666c395aa55ff71171915760401f39484c470946f39560a56bd67f9a34fe21f6ae8e6e472158e15363516fbd74380b72d324df2d44f54d61798d120e4eefeb444fb8845208de108fdc35f0babfce49bd3150c5f49d38b046350534196041275e4d8b9c97a5d3f50f4c6d61c8f4caf48)\n\n\n\n\n**Snap to grid**\npublished by @robikscube\nhttps://www.kaggle.com/robikscube/indoor-navigation-snap-to-grid-post-processing\nThe idea is to replace the predicted coordinates with the nearest waypoint of the training data. There are two positive effects:\n  1) pushing the predicted waypoint into the hallway when it is inside an obstacle.\n  2) enables us to predict the same point as the training data.\nAs you all know, there are a lot of overlapping waypoints between the training and the test data. So, selecting the coordinate from train waypoints is very effective. Since our team applied iterative post processing, we set threshold 6 in the first half post-processing parts to increase probability of snapping, and threshold 1.5 in the latter half to avoid strange snapping of already precise waypoints.\n\n**Parallel movement by Snap to grid**\nAs mentioned above, Snap to grid snaps predictions to the nearest point. However, the shape of the snapped path will not keep its original shape, especially most of the path is predicted to be inside a wall because snap direction is unpredictable. Therefore, before doing the snap to grid, we consider parallel moving of path to direction which is likely to be snapped while maintaining original shape. We calculate the difference between the original coordinates and the coordinates after Snap to Grid, and regard the average values of difference dx and dy as the direction in which Snap is likely to occur. By adding dx and dy to the original coordinates, we were able to move the path parallelly to the direction that is likely to be snapped while maintaining the original shape of the path.\n\n\n**Cost Minimization**\npublished by @saitodevel01\nhttps://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization\n \nThe idea is to minimize the difference between the distance to the predicted point and the distance calculated by the sensor data.\n \n⊿X^ is calculated by compute_rel_positions in compute_f. Though the host's prediction of relative coordinates with compute_rel_positions function is reasonably accurate, its accuracy depends on the quality of raw sensor data. If you look carefully at the relative coordinates, you will notice that the path of relative coordinates has been scaled up and that there is a bias in the direction of rotation compared to the original path. Let’s check the precision of this function.\n\n\n![](https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture3.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20220616T150947Z&X-Goog-Expires=345599&X-Goog-SignedHeaders=host&X-Goog-Signature=07cf865742e7f372aab6e278bc1b695315606c4104818bd9c1b04052c23d464e341c17ee6a361e5b39c893c3a67a438b25c178a6d674de09bf5d0d439df37076006e13f53db2ded49a9e2ed5da4d9f137d75bb6d18feee313799f4909651e27ccb850d8e6fe1a3cd850247a7cf71409c1b7692326aa7da8b3def2c683bae7cc63702fc2718ae8cbfc26faa80748c8be2df222df13da5736ee061996cd459979c638f5e4d907587124d00a7719f2ca18ca3e47518f408c8aec2434fc5135c6deb33558a5a6337b738ec1918076e54d96786a4e9b6acc09fa2905753d4d285404f6b42930196db37583d2dc63884a57ab73879718efa85b43501731c627ba2deea)\n\n\n\nBlue line indicates training path, and red line is relative positions. It is clear that relative position is scaled up and rotated compared to red line, so we need to calibrate compute_f function to get a good relative position. Compute_f function uses sensor data, so we try to clean up raw sensor data by Kalman Filter. Using Kalman filtered sensor data produces nice relative positions. We can get similar scale relative positions similar to the original path, which leads to a good score by cost minimization.\n\n\n\n![](https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture4.png?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20220616T151012Z&X-Goog-Expires=345599&X-Goog-SignedHeaders=host&X-Goog-Signature=198c723ffdd4fdb1de79a1deacda24fa8606b2d5ba0eb809ba0a90e4b22fdd237c1ae08f1b3bb261f2713e472f88c21cdc8a6250a53392398ba9b3702b50708dd602e4a2ad1328d5dac30e6e4d3eca7fe90a5ce7a269dece0ed30fcc4e3d94c78831586c666f158a8d211ea60d15f0c2c7016e93acd0f7b8ac79d30fb91727bed9f4383849659a3a84110274d0e7b230ff52c63d5eb9d0794af0ea31ba8ead3e7106d92142c9a46ee2682cee211dd17afb948f8432a143995067e6c06d4f06c0471ef8b529c8f84ccac431deb10d62cd58ba0bdcbe09ef66ba4b4d5177dd3b94f95a9ff792c1b9d39e967eba1a1e5efcb0078fbe34050325ab7c908e1f04d40b)\n\n\nNext, we tackle with the rotation problems. Here is an example.\n\n\n\n\n![](https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture5.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20220616T151036Z&X-Goog-Expires=345599&X-Goog-SignedHeaders=host&X-Goog-Signature=32e5d5e7c36cc9354a4c68341ac085a58c5b85953efe3e4980bcca888b514d74dbc114d916b4080ce7e258f55f8d9a5e57d495bed186d4cf0c7c0540528b245d4fa9caf36bf7347e5cdf18f7d05afb436a3aaffddef9c4f1edda32d0757096fdba1113e985423fa93dc314cfa078c600691740dbefcb38b3f9e0469b7ed87e41a04263dde99c7c098fb4b464001796ecf1587edaede52fa2a7019f00b5e74ff2cdaf14a7100d7b9d8a49ab38120f87eb83d2389d62aaba0856738e5b170f45e34df8d584b2263179b1129bb477db336be58fdc7ef1954c84e56f61663dde341d1bcc1008ab96b05f2bb522967f2e4963696d1934ec17160aa51b5892a78df5dd)\n\n\n\n\nIt is obvious that the red lines are rotated some degree. We try to solve these rotation problems with minimization of Euclid distance. \n\nLet the predicted coordinates in a certain path after post-processing be A1, A2, ... An in order from the start point, and the predicted coordinates by compute_rel_positions be B1, B2 ... Bn.If there is no rotation, the movement amount delta_a of A1 → A2 and the movement amount delta_b of B1 → B2 should be roughly the same. However, as you can see from the visualization, compute_rel_positions will rotate by some degrees, so delta_a and delta_b do not match. Therefore, we tried brute force search for the rotation angle θ of delta_b which minimizes the difference between delta_a and delta_b.\nThe loss function is the sum of the distances between delta_a and delta_b.\n\nloss (θ) = || (A2-A1) – (B2-B1) R (θ) || ^ 2 + || (A3-A2)-(B3-B2) R (θ) || ^ 2 +… + || (An-An-1) – (Bn-Bn-1) R (θ) || ^ 2\n\nWhere R (θ) is a rotation matrix that rotates the coordinates by θ degrees. For example, (B3 – B2) R (θ) represents B2 → B3 rotated by θ. Also, || X – Y || ^ 2 represents the square of the difference between the movements X and Y.In other words, the loss function represents the difference between the \"current movement amount\" and “the movement amount calculated by compute_rel_positions rotated by θ”.The smaller the value of the loss function, the closer the angle of B is to A.\n\n\n\n![](https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture6.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20220616T151051Z&X-Goog-Expires=345599&X-Goog-SignedHeaders=host&X-Goog-Signature=7fdaab9ded4bacb795239e4200aa3bbeff897865c5177983910adadc6d812262c9006766583e93310dec882e950b3199632f76b9fba5d63697bb806f39716ffe17f23c78384d4adca42b672263802c1835b8bcfd51fa95931b9765dc13fe8a2477f3ce957016696f3e01353f1612b41ec544fe220117fac3d302ee656737966f4460ccf423eb6b0f0b0e4097596ddf75711554448bca8e8b3fb1c544d011b15253bc98c97e04fa5388f39f1a3b7469007cd73951ff74acd0b8faffc723e2345c76abcd168b515699380cc3bb1dcfc25a03bf64373a926d1eaf7ee9f498df786848f7461dff5b7d0c079a4d36f871e5f36827192a0796705b4b4f57453658c412)\n\nWe modified the submission path by compute_rel_positions in the post process pipeline, and modified the rotation angle of compute_rel_positions by submission at the end of the pipeline. As we repeated the post process pipeline, both submission path and compute_rel_positions enhanced each other.\n \n\n \nThe difference between original and modified compute_f are very clear.\n\n\n![](https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture7.png?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20220616T151109Z&X-Goog-Expires=345599&X-Goog-SignedHeaders=host&X-Goog-Signature=3df41bfcad96849247ca344feda8468daa5b7569fb560d1df7fea21504685437ac7dcd56a8393b251b61e4e2af32ab8824ab0df56476e90965f8f0286ff0e8fd0f0bb775da3b46fc93a9495030c44a5542679ff0a2cb7e14106b92fdeb36cf0e340c75a3b3960b4cfbf3f68b1ab938945060bf9f86d17a13dac01bdba56e85a6d7dcd7dcc1b86ebaa7bd3e016a0f279d60d3f55189eba80a48e00afd71cb04d07a1284508fce1b92372624c4e9099298f24dfcf431d42da59fb98a503b726e75c952281089d4715e3289003cc3209400cd95d5406556cade8d14be1067366f8427b53ac3e7e4054432bf67b7c6f084de8047513bfc58ed94804afb0b5caf3203)\n\n![](https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture8.png?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20220616T151123Z&X-Goog-Expires=345599&X-Goog-SignedHeaders=host&X-Goog-Signature=b073cb1f4daffe9761010eefa4cf2387650b188f410fec31029d00f163b6575e7e0f7bc3e5cf2987c234e355a14bebe892192afeac637262d15ef934ce594085df4960184a9ad7e4c988cc119708c5349192b808dc1d4b47d0522bb6b50a2647626ebe66e17db70b71f3222ea2774f4a24a29ddb1f97710d93f0ff99ba57ddb4e2201a71ef755aad92720bc8cf7e2d42b62e25abf985e2b5498e2193c1b9fcc00b4cdd8f4c563d0687f7694cb0e49ad411a7e915880aa435ef4dd03e38040782aef7f8ed90777e16495ae8407011483427fe00325de9f5ef146be85202f417a1c6ab5241afa9ab3555e842b31694f2b299a210f2df5f172ff847a903ff76ac82)\n\n\n In addition, cost minimization has two important parameters, α and β. The higher the α, the more emphasize the current prediction, and the higher the β, the more emphasize the sensor data. Since we use Cost Minimization many times in our pipeline, β was attenuated with each successive use so that Cost Minimization would not be too influenced by the sensor data.\n\n\n\n\n**Push to Hallway**\nIf predictions are located in an obstacle, we need to push them to the nearest waypoint. But when we just increase the threshold of Snap to Grid, waypoints are changed to strange waypoints since there are many hallways which don’t have waypoints. So, we tried to push the waypoints to the nearest hallway even if there were no waypoints.\nEach pixel of the floor image was used to determine whether predictions are located in the hallway or an obstacle, and moved waypoints located in the obstacle to the nearest coordinates in the hallway. This has the great advantage of being able to modify the waypoints in obstacles which is far from training waypoints.\n\n\n![](https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture9.png?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20220616T151136Z&X-Goog-Expires=345599&X-Goog-SignedHeaders=host&X-Goog-Signature=23a9f942b3a842db1df2d1ab7a051ffca9450987bcc796b7c492e8749d40b8845c41f32842723abe09e4ca53a2fa53a5894a235f1108fb4f9a175653249bbda8ef80d92406c954508358c822c29dbfd77575541191e5082c2ad71a1fd233f98a6ad11f396e1e38f60519adc63520727f9e6c85e7af77707f72273e3fc0f11eebf0e82b7d31bf3c40b38ce95ccf4fc21813e23445ad1388fcd6b7fd5a7b4cf98a6d87a5e8b125e11b6fb2b216df689a1fed4e0937f22841ba5e3aa3c7ebab450de19305b2bd3ebcf2ddd12ba451bcf74eb89b50df135176f115d7133d915af0f79472a286b609ed3d542616bc5f471da8abfeac66b146f1f29345ae033550c1b1)\n\n\n**Apply Start and End Points leakage**\npublished by @tomooinubushi\nhttps://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage\n\nThis was beyond our analytical capabilities, so we only corrected the start and end waypoints as same as the public notebook.\n\n**What didn’t work**\nTime series RNN.\nInput all WiFi and sensor data to NN.\neach model for each site.\nPre training of all data and Transfer Learning.\nDijkstra, Warshall-Froyd, Band First Search as shortest path solving \nMap Matching\nCurve to Curve\nGrid Generation\n\n\n\n\nThank you for reading our solution. \nAny opinions are welcome!",
      "votes": null
    },
    {
      "id": "1317652",
      "postDate": "05/21/2021 15:20:40",
      "content": "<p><a href=\"https://www.kaggle.com/cocoinit23\" target=\"_blank\">@cocoinit23</a>, not only is this a really elegant slution, but I also love the kid's sketch map of the USA that appears as the first red, white and blue Figure in your Cost Minimization paragraph.</p>",
      "rawMarkdown": "cocoinit23, not only is this a really elegant slution, but I also love the kid's sketch map of the USA that appears as the first red, white and blue Figure in your Cost Minimization paragraph.",
      "votes": null
    },
    {
      "id": "1317664",
      "postDate": "05/21/2021 15:26:47",
      "content": "<p>Congrats on achieving!!! Thanks for sharing 👍</p>",
      "rawMarkdown": "Congrats on achieving!!! Thanks for sharing 👍",
      "votes": null
    },
    {
      "id": "1320369",
      "postDate": "05/24/2021 03:14:42",
      "content": "<p>Congratz for the achievement!<br>\nWould you plan to release your code, would love to learn from that!</p>",
      "rawMarkdown": "Congratz for the achievement!\nWould you plan to release your code, would love to learn from that!",
      "votes": null
    },
    {
      "id": "1321416",
      "postDate": "05/24/2021 17:08:21",
      "content": "<p>thank you for sharing!</p>",
      "rawMarkdown": "thank you for sharing!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1317652,
      "author_name": "jbomitchell",
      "author_url": "",
      "post_date": "05/21/2021 15:20:40",
      "content": "<p><a href=\"https://www.kaggle.com/cocoinit23\" target=\"_blank\">@cocoinit23</a>, not only is this a really elegant slution, but I also love the kid's sketch map of the USA that appears as the first red, white and blue Figure in your Cost Minimization paragraph.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1317664,
      "author_name": "ankitp013",
      "author_url": "",
      "post_date": "05/21/2021 15:26:47",
      "content": "<p>Congrats on achieving!!! Thanks for sharing 👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1320369,
      "author_name": "alexlwh",
      "author_url": "",
      "post_date": "05/24/2021 03:14:42",
      "content": "<p>Congratz for the achievement!<br>\nWould you plan to release your code, would love to learn from that!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1321416,
      "author_name": "harshitsati",
      "author_url": "",
      "post_date": "05/24/2021 17:08:21",
      "content": "<p>thank you for sharing!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1317353": "Thank the host for such an interesting competition. And also thank my great teammates for working hard throughout the competition ( @ryotayoshinobu, @tubotubo, @columbia2131 ).\nThe key factor of this competition is post processing based on algorithms rather than modeling. Also, all of us like competitive programming, so there is no other team’s name but “Algorithm is All You Need” :)\n \nHere is our solution.\n\n\n**Quick summary**\n\n\n![](https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture1.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20220616T150747Z&X-Goog-Expires=345599&X-Goog-SignedHeaders=host&X-Goog-Signature=7de5d6ffefbd8e3fa8770da60dbacdec1604c7c080f3bf63a729f7466c446bbcc72b2ea6425bd349b5732571440fe2a7856a80873dfdaadbec80d9f1104409bd2940ec32abeae4d516a6f84c5dfa2c25308abdd160a1c199f3c2f693ac7dd3217adf9b4fdb12fdaa8d2369d081338f9cd12552af7a0e0f1f28fc2bdef89b4e43b6163cfb77d2aec9125af12c27874e079d5b99dd1625b6455118cd93c3a0e6c071d0cc2d8f7c7a33c589b0183226cb191a773eb547a23e6f906e56432765827fd0f5156eab0daaae6c528329febf77d2797ad2b128ccbe18ad9df23406520fe921f79ccc2e85d6dd3492d8637b73d74086c4134d4c77608a1a7d61878a38c4bb)\n\nThis competition required a good amount of pre and post processing. Our final prediction was a blend of 3rd stage models trained with iterative pseudo labeling. Repeating the post processing was also important for our solution.\n\n**Simple 100% Accurate Floor model**\n\n- Consider RSSI as the number of occurrences of a BSSID.\n- TF-IDF vectorization.\n- KNN for each site.\n\nOur KNN model performed 100% accuracy for both public and private dataset.\n\n\n**Waypoint Model**\n\n**Preprocess**\n-   \tLinear interpolation of waypoints based on WiFi timestamp.\n-   \tFix malformed txt data @higepon ‘s notebook \nhttps://www.kaggle.com/higepon/how-to-fix-malformed-train-test-data\n-   \tRemove WiFi information which exists only in training datasets.\n-   \tApplying Kalman Filter to sensor data for getting a precise result of compute_f.\n\n\n\n**Cross Validation**\n-   \tGroupKFold of path\n-   \t5 CV but n_splits=15 to get many combinations of path groups to enhance random seed averaging.\n\n\n\n**1st stage training**\n-   \tLSTM based on @Kouki ‘s notebook \nhttps://www.kaggle.com/kokitanisaka/lstm-by-keras-with-unified-wi-fi-feats.\nHyper parameters were tuned with Optuna.\n-   \tTransformer-like Convolutional Encoder.\nEmbedding BSSID and RSSI, then Positional Encoding are added.\nNo Feed Forward Network.\nBatch Normalization instead of Layer Normalization.\nMultiplying attention instead of Adding.\nQuery of MultiHeadAttention is BSSID embeddings, Key and Value are RSSI embeddings.\nEncoding process is like [Embedding -> Conv1d -> Multiply Attention -> Conv1d -> Multiply Attention -> … -> Dense]\n-   \t10 random seeds to learn multiple combinations of path groups.\n\n\n\n**2nd stage training**\n-   \tConvolutional stacking to learn correlation between different models or different random seeds, which is equal to learning multiple path combinations.\n-   \tLightGBM to learn time series and relative position information.\nAdditional input features are:\n1. n predictions before and after\n2. Difference between n predictions\n3. Rate of change from n predictions\n4. Difference between the previous and next n predictions\n5. Moving average\n6. Moving variance\n7. Relative position (compute_f, compute_rel_position)\n8. Cumulative sum of relative positions\n9. Difference between predicted values and cumulative sum of relative positions\n10. Aggregate features for each path (mean,max,min,median,std,sum)\n \n\n\n**3rd stage training**\n-   \tHill Climbing to get weights which minimize loss.\n-   \tRidge regression to suppress overfitting.\n \n \n**Post process**\nPost processing is very important to push up scores significantly, and also a fun part for competitive programmers. Here we describe detailed steps. Summary is described below figure.\n\n\n![](https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture2.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20220616T150904Z&X-Goog-Expires=345599&X-Goog-SignedHeaders=host&X-Goog-Signature=937a2b44541155f08662c1318ddbc7617da25e94aaf7e15e38c04ebadbd033350968cc93dbf5fc6ab8443f17d3d9d343ed5363cc3099a77fff710b2ea5d3133b82fc5af18813b6fd4d90cb2940b3036696ad144c3697554209973e0b1b355f164623746548bda9d05330b37b5aed9fdd2ee6a6d89adb9766b7818ea89cd2afddbaad16538844dcf3cc62db42b062c18f6666c395aa55ff71171915760401f39484c470946f39560a56bd67f9a34fe21f6ae8e6e472158e15363516fbd74380b72d324df2d44f54d61798d120e4eefeb444fb8845208de108fdc35f0babfce49bd3150c5f49d38b046350534196041275e4d8b9c97a5d3f50f4c6d61c8f4caf48)\n\n\n\n\n**Snap to grid**\npublished by @robikscube\nhttps://www.kaggle.com/robikscube/indoor-navigation-snap-to-grid-post-processing\nThe idea is to replace the predicted coordinates with the nearest waypoint of the training data. There are two positive effects:\n  1) pushing the predicted waypoint into the hallway when it is inside an obstacle.\n  2) enables us to predict the same point as the training data.\nAs you all know, there are a lot of overlapping waypoints between the training and the test data. So, selecting the coordinate from train waypoints is very effective. Since our team applied iterative post processing, we set threshold 6 in the first half post-processing parts to increase probability of snapping, and threshold 1.5 in the latter half to avoid strange snapping of already precise waypoints.\n\n**Parallel movement by Snap to grid**\nAs mentioned above, Snap to grid snaps predictions to the nearest point. However, the shape of the snapped path will not keep its original shape, especially most of the path is predicted to be inside a wall because snap direction is unpredictable. Therefore, before doing the snap to grid, we consider parallel moving of path to direction which is likely to be snapped while maintaining original shape. We calculate the difference between the original coordinates and the coordinates after Snap to Grid, and regard the average values of difference dx and dy as the direction in which Snap is likely to occur. By adding dx and dy to the original coordinates, we were able to move the path parallelly to the direction that is likely to be snapped while maintaining the original shape of the path.\n\n\n**Cost Minimization**\npublished by @saitodevel01\nhttps://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization\n \nThe idea is to minimize the difference between the distance to the predicted point and the distance calculated by the sensor data.\n \n⊿X^ is calculated by compute_rel_positions in compute_f. Though the host's prediction of relative coordinates with compute_rel_positions function is reasonably accurate, its accuracy depends on the quality of raw sensor data. If you look carefully at the relative coordinates, you will notice that the path of relative coordinates has been scaled up and that there is a bias in the direction of rotation compared to the original path. Let’s check the precision of this function.\n\n\n![](https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture3.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20220616T150947Z&X-Goog-Expires=345599&X-Goog-SignedHeaders=host&X-Goog-Signature=07cf865742e7f372aab6e278bc1b695315606c4104818bd9c1b04052c23d464e341c17ee6a361e5b39c893c3a67a438b25c178a6d674de09bf5d0d439df37076006e13f53db2ded49a9e2ed5da4d9f137d75bb6d18feee313799f4909651e27ccb850d8e6fe1a3cd850247a7cf71409c1b7692326aa7da8b3def2c683bae7cc63702fc2718ae8cbfc26faa80748c8be2df222df13da5736ee061996cd459979c638f5e4d907587124d00a7719f2ca18ca3e47518f408c8aec2434fc5135c6deb33558a5a6337b738ec1918076e54d96786a4e9b6acc09fa2905753d4d285404f6b42930196db37583d2dc63884a57ab73879718efa85b43501731c627ba2deea)\n\n\n\nBlue line indicates training path, and red line is relative positions. It is clear that relative position is scaled up and rotated compared to red line, so we need to calibrate compute_f function to get a good relative position. Compute_f function uses sensor data, so we try to clean up raw sensor data by Kalman Filter. Using Kalman filtered sensor data produces nice relative positions. We can get similar scale relative positions similar to the original path, which leads to a good score by cost minimization.\n\n\n\n![](https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture4.png?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20220616T151012Z&X-Goog-Expires=345599&X-Goog-SignedHeaders=host&X-Goog-Signature=198c723ffdd4fdb1de79a1deacda24fa8606b2d5ba0eb809ba0a90e4b22fdd237c1ae08f1b3bb261f2713e472f88c21cdc8a6250a53392398ba9b3702b50708dd602e4a2ad1328d5dac30e6e4d3eca7fe90a5ce7a269dece0ed30fcc4e3d94c78831586c666f158a8d211ea60d15f0c2c7016e93acd0f7b8ac79d30fb91727bed9f4383849659a3a84110274d0e7b230ff52c63d5eb9d0794af0ea31ba8ead3e7106d92142c9a46ee2682cee211dd17afb948f8432a143995067e6c06d4f06c0471ef8b529c8f84ccac431deb10d62cd58ba0bdcbe09ef66ba4b4d5177dd3b94f95a9ff792c1b9d39e967eba1a1e5efcb0078fbe34050325ab7c908e1f04d40b)\n\n\nNext, we tackle with the rotation problems. Here is an example.\n\n\n\n\n![](https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture5.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20220616T151036Z&X-Goog-Expires=345599&X-Goog-SignedHeaders=host&X-Goog-Signature=32e5d5e7c36cc9354a4c68341ac085a58c5b85953efe3e4980bcca888b514d74dbc114d916b4080ce7e258f55f8d9a5e57d495bed186d4cf0c7c0540528b245d4fa9caf36bf7347e5cdf18f7d05afb436a3aaffddef9c4f1edda32d0757096fdba1113e985423fa93dc314cfa078c600691740dbefcb38b3f9e0469b7ed87e41a04263dde99c7c098fb4b464001796ecf1587edaede52fa2a7019f00b5e74ff2cdaf14a7100d7b9d8a49ab38120f87eb83d2389d62aaba0856738e5b170f45e34df8d584b2263179b1129bb477db336be58fdc7ef1954c84e56f61663dde341d1bcc1008ab96b05f2bb522967f2e4963696d1934ec17160aa51b5892a78df5dd)\n\n\n\n\nIt is obvious that the red lines are rotated some degree. We try to solve these rotation problems with minimization of Euclid distance. \n\nLet the predicted coordinates in a certain path after post-processing be A1, A2, ... An in order from the start point, and the predicted coordinates by compute_rel_positions be B1, B2 ... Bn.If there is no rotation, the movement amount delta_a of A1 → A2 and the movement amount delta_b of B1 → B2 should be roughly the same. However, as you can see from the visualization, compute_rel_positions will rotate by some degrees, so delta_a and delta_b do not match. Therefore, we tried brute force search for the rotation angle θ of delta_b which minimizes the difference between delta_a and delta_b.\nThe loss function is the sum of the distances between delta_a and delta_b.\n\nloss (θ) = || (A2-A1) – (B2-B1) R (θ) || ^ 2 + || (A3-A2)-(B3-B2) R (θ) || ^ 2 +… + || (An-An-1) – (Bn-Bn-1) R (θ) || ^ 2\n\nWhere R (θ) is a rotation matrix that rotates the coordinates by θ degrees. For example, (B3 – B2) R (θ) represents B2 → B3 rotated by θ. Also, || X – Y || ^ 2 represents the square of the difference between the movements X and Y.In other words, the loss function represents the difference between the \"current movement amount\" and “the movement amount calculated by compute_rel_positions rotated by θ”.The smaller the value of the loss function, the closer the angle of B is to A.\n\n\n\n![](https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture6.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20220616T151051Z&X-Goog-Expires=345599&X-Goog-SignedHeaders=host&X-Goog-Signature=7fdaab9ded4bacb795239e4200aa3bbeff897865c5177983910adadc6d812262c9006766583e93310dec882e950b3199632f76b9fba5d63697bb806f39716ffe17f23c78384d4adca42b672263802c1835b8bcfd51fa95931b9765dc13fe8a2477f3ce957016696f3e01353f1612b41ec544fe220117fac3d302ee656737966f4460ccf423eb6b0f0b0e4097596ddf75711554448bca8e8b3fb1c544d011b15253bc98c97e04fa5388f39f1a3b7469007cd73951ff74acd0b8faffc723e2345c76abcd168b515699380cc3bb1dcfc25a03bf64373a926d1eaf7ee9f498df786848f7461dff5b7d0c079a4d36f871e5f36827192a0796705b4b4f57453658c412)\n\nWe modified the submission path by compute_rel_positions in the post process pipeline, and modified the rotation angle of compute_rel_positions by submission at the end of the pipeline. As we repeated the post process pipeline, both submission path and compute_rel_positions enhanced each other.\n \n\n \nThe difference between original and modified compute_f are very clear.\n\n\n![](https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture7.png?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20220616T151109Z&X-Goog-Expires=345599&X-Goog-SignedHeaders=host&X-Goog-Signature=3df41bfcad96849247ca344feda8468daa5b7569fb560d1df7fea21504685437ac7dcd56a8393b251b61e4e2af32ab8824ab0df56476e90965f8f0286ff0e8fd0f0bb775da3b46fc93a9495030c44a5542679ff0a2cb7e14106b92fdeb36cf0e340c75a3b3960b4cfbf3f68b1ab938945060bf9f86d17a13dac01bdba56e85a6d7dcd7dcc1b86ebaa7bd3e016a0f279d60d3f55189eba80a48e00afd71cb04d07a1284508fce1b92372624c4e9099298f24dfcf431d42da59fb98a503b726e75c952281089d4715e3289003cc3209400cd95d5406556cade8d14be1067366f8427b53ac3e7e4054432bf67b7c6f084de8047513bfc58ed94804afb0b5caf3203)\n\n![](https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture8.png?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20220616T151123Z&X-Goog-Expires=345599&X-Goog-SignedHeaders=host&X-Goog-Signature=b073cb1f4daffe9761010eefa4cf2387650b188f410fec31029d00f163b6575e7e0f7bc3e5cf2987c234e355a14bebe892192afeac637262d15ef934ce594085df4960184a9ad7e4c988cc119708c5349192b808dc1d4b47d0522bb6b50a2647626ebe66e17db70b71f3222ea2774f4a24a29ddb1f97710d93f0ff99ba57ddb4e2201a71ef755aad92720bc8cf7e2d42b62e25abf985e2b5498e2193c1b9fcc00b4cdd8f4c563d0687f7694cb0e49ad411a7e915880aa435ef4dd03e38040782aef7f8ed90777e16495ae8407011483427fe00325de9f5ef146be85202f417a1c6ab5241afa9ab3555e842b31694f2b299a210f2df5f172ff847a903ff76ac82)\n\n\n In addition, cost minimization has two important parameters, α and β. The higher the α, the more emphasize the current prediction, and the higher the β, the more emphasize the sensor data. Since we use Cost Minimization many times in our pipeline, β was attenuated with each successive use so that Cost Minimization would not be too influenced by the sensor data.\n\n\n\n\n**Push to Hallway**\nIf predictions are located in an obstacle, we need to push them to the nearest waypoint. But when we just increase the threshold of Snap to Grid, waypoints are changed to strange waypoints since there are many hallways which don’t have waypoints. So, we tried to push the waypoints to the nearest hallway even if there were no waypoints.\nEach pixel of the floor image was used to determine whether predictions are located in the hallway or an obstacle, and moved waypoints located in the obstacle to the nearest coordinates in the hallway. This has the great advantage of being able to modify the waypoints in obstacles which is far from training waypoints.\n\n\n![](https://storage.googleapis.com/kagglesdsdata/datasets/1357497/2256047/Picture9.png?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20220616%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20220616T151136Z&X-Goog-Expires=345599&X-Goog-SignedHeaders=host&X-Goog-Signature=23a9f942b3a842db1df2d1ab7a051ffca9450987bcc796b7c492e8749d40b8845c41f32842723abe09e4ca53a2fa53a5894a235f1108fb4f9a175653249bbda8ef80d92406c954508358c822c29dbfd77575541191e5082c2ad71a1fd233f98a6ad11f396e1e38f60519adc63520727f9e6c85e7af77707f72273e3fc0f11eebf0e82b7d31bf3c40b38ce95ccf4fc21813e23445ad1388fcd6b7fd5a7b4cf98a6d87a5e8b125e11b6fb2b216df689a1fed4e0937f22841ba5e3aa3c7ebab450de19305b2bd3ebcf2ddd12ba451bcf74eb89b50df135176f115d7133d915af0f79472a286b609ed3d542616bc5f471da8abfeac66b146f1f29345ae033550c1b1)\n\n\n**Apply Start and End Points leakage**\npublished by @tomooinubushi\nhttps://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage\n\nThis was beyond our analytical capabilities, so we only corrected the start and end waypoints as same as the public notebook.\n\n**What didn’t work**\nTime series RNN.\nInput all WiFi and sensor data to NN.\neach model for each site.\nPre training of all data and Transfer Learning.\nDijkstra, Warshall-Froyd, Band First Search as shortest path solving \nMap Matching\nCurve to Curve\nGrid Generation\n\n\n\n\nThank you for reading our solution. \nAny opinions are welcome!",
    "1317652": "cocoinit23, not only is this a really elegant slution, but I also love the kid's sketch map of the USA that appears as the first red, white and blue Figure in your Cost Minimization paragraph.",
    "1317664": "Congrats on achieving!!! Thanks for sharing 👍",
    "1320369": "Congratz for the achievement!\nWould you plan to release your code, would love to learn from that!",
    "1321416": "thank you for sharing!"
  },
  "source": "meta"
}