{
  "id": 240196,
  "title": "12th place solution",
  "url": "/competitions/indoor-location-navigation/writeups/hugues-12th-place-solution",
  "author_name": "",
  "post_date": "2021-05-18T20:55:54.747Z",
  "votes": 13,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Congratulations to all the winners. Even if I had no experience in the domain, I found this competition very interesting and decided to enter it very early. This allowed me to try different ideas. Let me present the one that I used in my final submission.</p>\n<p><strong>1 - Pre-processing</strong><br>\nI used only wifi signals for positions prediction. Beacon or magnetic data didn't help.<br>\nI grouped all wifi rows by block, but I reallocated wifi rows based on \"last seen timestamp\" to the wifi block that was the closest with respect to timestamp.</p>\n<p><strong>2 - Floor prediction and Wifi-based position prediction</strong><br>\nIn a first step, I used LGBM models for floor prediction (1 model by building) and simple 2 layers feed forward NNs for (x, y) positions (1 model for each floor). The performance was OK.<br>\nThen, I switch to a rather radical approach that was providing better results:<br>\nI computed for each test point the cosine of the angle of wifi fingerprints with respect to each training point. That is: the scal prod divided by the norm of the 2 vectors.<br>\nUsing the cosine instead of the actual scal prod is important because the intensity of the wifi signals varies a lot for different positions.<br>\nAs a result, a perfect match would return a cosine of 1, and the value decreases as the closest training point \"match\" is farther away (actually, there is also a bit of post-processing to discard outlayers).<br>\nFor each test point, the best training point \"match\" is computed FOR EACH FLOOR.<br>\nThe floor prediction is then performed by checking the evolution of the cosine along the path: the floor for which the cosine is the highest more frequently is selected as the predicted floor.</p>\n<p>Using this approach (which is a kind of k-NN), there is no model (no parameter to train), and no CV…<br>\nActually, following the simple idea presented in the thread: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization</a>, I submitted an edited version of one of my final submissions and was happy to see that my floor prediction was 100% accurate both for public and private data:</p>\n<table>\n<thead>\n<tr>\n<th>Submission and Description</th>\n<th>Private Score</th>\n<th>Public Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>submission_df_leak_start_end - REFINED_PATHS-FullPath-SE-NoENF-LS_submission-2021-05-17_07-58-29 - score2.692.csv</td>\n<td>3.25303</td>\n<td>2.69255</td>\n</tr>\n<tr>\n<td>submission_INCREMENT_2021-05-18_17-42-06.csv</td>\n<td>18.25303</td>\n<td>17.69255</td>\n</tr>\n</tbody>\n</table>\n<p><strong>3 - Introducing accelerometer data</strong><br>\nWith the approch discussed in section 2, the position prediction is still very noisy. It is necessary to exploit the accelerometer data.<br>\nAt this stage, I didn't start from the raw data but I used instead the library provided by the organisers (\"compute_step_positions\").<br>\nFirst, I used my own post-processing (using some local averaging along paths, combining initial (x, y) predictions and step positions using a moving window). However, Saito's notebook was much more efficient. Also, the weights in the cost minimization expression could use the cosine values to reflect points for which one is more confident.</p>\n<p><strong>4 - Improving paths by local search</strong><br>\nThe output of section 3 is a list of (x, y) positions that define paths.<br>\nI defined another cost function that combined 3 values for each path:</p>\n<ul>\n<li>error for (x, y) positions (current positions vs \"starting\" positions, for each point on the path)</li>\n<li>error for length of each segment in the path (ratio between length based on start/end positions, and associated step_positions length)</li>\n<li>error for heading between consecutive segments (again: variation in heading based on (x, y) positions vs variations in heading based on accelerometer data)</li>\n</ul>\n<p>Then, a population-based search was performed using some \"mutation\" operators to update the path and capture constraints to place points in corridors.<br>\nAgain, no CV framework was used here, but improvements in the local search based on this cost function appeared highly correlated with LB scores.</p>\n<p><strong>5 - Post-processing</strong><br>\nI didn't explore that part that much. I just lazily invoked the snap-to-grid (<a href=\"url\" target=\"_blank\">https://www.kaggle.com/dragonzhang/3-3-g6-indoor-navigation-snap-to-grid</a>) and leakage (<a href=\"url\" target=\"_blank\">https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage</a>) notebooks like many of us :-).<br>\nThese notebooks were regularly improving my solutions by a value of around 0.3.</p>\n<p><strong>Final comments</strong><br>\nOne direction I wanted to investigate was about correcting step_positions returned by the python library provided by the competition host. Indeed, doing some statistics with training data, one could notice significant discrepancies between segments based on actual training positions and corresponding step_positions (roughly, with a factor in the range: 0.5 to 1.5, depending on paths).</p>\n<p>I think the main benefit of the approach presented above is that it does not rely too much on training points (except for final snap to grid). I believe this makes it more applicable to real life applications for which positions to predict don't match known positions… However, it does not take advantage enough of the setup of this competition for which a large fraction of test positions matches training positions!</p>\n<p><strong>References:</strong><br>\nI want to thank the authors of the following 4 notebooks that helped me a lot:<br>\n1) <a href=\"url\" target=\"_blank\">https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization</a><br>\nLike many of us, I also used Saito's code in my approach. At the beginning, this notebook was not improving my solution that much, but runtime was much better!<br>\n2) <a href=\"url\" target=\"_blank\">https://www.kaggle.com/rafaelcartenet/scaled-floors-geojsons-new-dataset</a><br>\nI wasn't familiar with the shapely library. This notebook provided me with all the code for supporting the local search approach that I implemented. Many thanks to Rafael Cartenet!<br>\n3) <a href=\"url\" target=\"_blank\">https://www.kaggle.com/dragonzhang/3-3-g6-indoor-navigation-snap-to-grid</a><br>\n4) <a href=\"url\" target=\"_blank\">https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage</a></p>",
  "messages": [
    {
      "id": "1313955",
      "postDate": "05/18/2021 20:24:49",
      "content": "<p>Congratulations to all the winners. Even if I had no experience in the domain, I found this competition very interesting and decided to enter it very early. This allowed me to try different ideas. Let me present the one that I used in my final submission.</p>\n<p><strong>1 - Pre-processing</strong><br>\nI used only wifi signals for positions prediction. Beacon or magnetic data didn't help.<br>\nI grouped all wifi rows by block, but I reallocated wifi rows based on \"last seen timestamp\" to the wifi block that was the closest with respect to timestamp.</p>\n<p><strong>2 - Floor prediction and Wifi-based position prediction</strong><br>\nIn a first step, I used LGBM models for floor prediction (1 model by building) and simple 2 layers feed forward NNs for (x, y) positions (1 model for each floor). The performance was OK.<br>\nThen, I switch to a rather radical approach that was providing better results:<br>\nI computed for each test point the cosine of the angle of wifi fingerprints with respect to each training point. That is: the scal prod divided by the norm of the 2 vectors.<br>\nUsing the cosine instead of the actual scal prod is important because the intensity of the wifi signals varies a lot for different positions.<br>\nAs a result, a perfect match would return a cosine of 1, and the value decreases as the closest training point \"match\" is farther away (actually, there is also a bit of post-processing to discard outlayers).<br>\nFor each test point, the best training point \"match\" is computed FOR EACH FLOOR.<br>\nThe floor prediction is then performed by checking the evolution of the cosine along the path: the floor for which the cosine is the highest more frequently is selected as the predicted floor.</p>\n<p>Using this approach (which is a kind of k-NN), there is no model (no parameter to train), and no CV…<br>\nActually, following the simple idea presented in the thread: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization</a>, I submitted an edited version of one of my final submissions and was happy to see that my floor prediction was 100% accurate both for public and private data:</p>\n<table>\n<thead>\n<tr>\n<th>Submission and Description</th>\n<th>Private Score</th>\n<th>Public Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>submission_df_leak_start_end - REFINED_PATHS-FullPath-SE-NoENF-LS_submission-2021-05-17_07-58-29 - score2.692.csv</td>\n<td>3.25303</td>\n<td>2.69255</td>\n</tr>\n<tr>\n<td>submission_INCREMENT_2021-05-18_17-42-06.csv</td>\n<td>18.25303</td>\n<td>17.69255</td>\n</tr>\n</tbody>\n</table>\n<p><strong>3 - Introducing accelerometer data</strong><br>\nWith the approch discussed in section 2, the position prediction is still very noisy. It is necessary to exploit the accelerometer data.<br>\nAt this stage, I didn't start from the raw data but I used instead the library provided by the organisers (\"compute_step_positions\").<br>\nFirst, I used my own post-processing (using some local averaging along paths, combining initial (x, y) predictions and step positions using a moving window). However, Saito's notebook was much more efficient. Also, the weights in the cost minimization expression could use the cosine values to reflect points for which one is more confident.</p>\n<p><strong>4 - Improving paths by local search</strong><br>\nThe output of section 3 is a list of (x, y) positions that define paths.<br>\nI defined another cost function that combined 3 values for each path:</p>\n<ul>\n<li>error for (x, y) positions (current positions vs \"starting\" positions, for each point on the path)</li>\n<li>error for length of each segment in the path (ratio between length based on start/end positions, and associated step_positions length)</li>\n<li>error for heading between consecutive segments (again: variation in heading based on (x, y) positions vs variations in heading based on accelerometer data)</li>\n</ul>\n<p>Then, a population-based search was performed using some \"mutation\" operators to update the path and capture constraints to place points in corridors.<br>\nAgain, no CV framework was used here, but improvements in the local search based on this cost function appeared highly correlated with LB scores.</p>\n<p><strong>5 - Post-processing</strong><br>\nI didn't explore that part that much. I just lazily invoked the snap-to-grid (<a href=\"url\" target=\"_blank\">https://www.kaggle.com/dragonzhang/3-3-g6-indoor-navigation-snap-to-grid</a>) and leakage (<a href=\"url\" target=\"_blank\">https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage</a>) notebooks like many of us :-).<br>\nThese notebooks were regularly improving my solutions by a value of around 0.3.</p>\n<p><strong>Final comments</strong><br>\nOne direction I wanted to investigate was about correcting step_positions returned by the python library provided by the competition host. Indeed, doing some statistics with training data, one could notice significant discrepancies between segments based on actual training positions and corresponding step_positions (roughly, with a factor in the range: 0.5 to 1.5, depending on paths).</p>\n<p>I think the main benefit of the approach presented above is that it does not rely too much on training points (except for final snap to grid). I believe this makes it more applicable to real life applications for which positions to predict don't match known positions… However, it does not take advantage enough of the setup of this competition for which a large fraction of test positions matches training positions!</p>\n<p><strong>References:</strong><br>\nI want to thank the authors of the following 4 notebooks that helped me a lot:<br>\n1) <a href=\"url\" target=\"_blank\">https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization</a><br>\nLike many of us, I also used Saito's code in my approach. At the beginning, this notebook was not improving my solution that much, but runtime was much better!<br>\n2) <a href=\"url\" target=\"_blank\">https://www.kaggle.com/rafaelcartenet/scaled-floors-geojsons-new-dataset</a><br>\nI wasn't familiar with the shapely library. This notebook provided me with all the code for supporting the local search approach that I implemented. Many thanks to Rafael Cartenet!<br>\n3) <a href=\"url\" target=\"_blank\">https://www.kaggle.com/dragonzhang/3-3-g6-indoor-navigation-snap-to-grid</a><br>\n4) <a href=\"url\" target=\"_blank\">https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage</a></p>",
      "rawMarkdown": "Congratulations to all the winners. Even if I had no experience in the domain, I found this competition very interesting and decided to enter it very early. This allowed me to try different ideas. Let me present the one that I used in my final submission.\n\n**1 - Pre-processing**\nI used only wifi signals for positions prediction. Beacon or magnetic data didn't help.\nI grouped all wifi rows by block, but I reallocated wifi rows based on \"last seen timestamp\" to the wifi block that was the closest with respect to timestamp.\n\n**2 - Floor prediction and Wifi-based position prediction**\nIn a first step, I used LGBM models for floor prediction (1 model by building) and simple 2 layers feed forward NNs for (x, y) positions (1 model for each floor). The performance was OK.\nThen, I switch to a rather radical approach that was providing better results:\nI computed for each test point the cosine of the angle of wifi fingerprints with respect to each training point. That is: the scal prod divided by the norm of the 2 vectors.\nUsing the cosine instead of the actual scal prod is important because the intensity of the wifi signals varies a lot for different positions.\nAs a result, a perfect match would return a cosine of 1, and the value decreases as the closest training point \"match\" is farther away (actually, there is also a bit of post-processing to discard outlayers).\nFor each test point, the best training point \"match\" is computed FOR EACH FLOOR.\nThe floor prediction is then performed by checking the evolution of the cosine along the path: the floor for which the cosine is the highest more frequently is selected as the predicted floor.\n\nUsing this approach (which is a kind of k-NN), there is no model (no parameter to train), and no CV...\nActually, following the simple idea presented in the thread: [https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization](url), I submitted an edited version of one of my final submissions and was happy to see that my floor prediction was 100% accurate both for public and private data:\n\n| Submission and Description | Private Score | Public Score |\n| --- | --- | --- |\n| submission_df_leak_start_end - REFINED_PATHS-FullPath-SE-NoENF-LS_submission-2021-05-17_07-58-29 - score2.692.csv | 3.25303 | 2.69255 |\n| submission_INCREMENT_2021-05-18_17-42-06.csv | 18.25303 | 17.69255 |\n\n**3 - Introducing accelerometer data**\nWith the approch discussed in section 2, the position prediction is still very noisy. It is necessary to exploit the accelerometer data.\nAt this stage, I didn't start from the raw data but I used instead the library provided by the organisers (\"compute_step_positions\").\nFirst, I used my own post-processing (using some local averaging along paths, combining initial (x, y) predictions and step positions using a moving window). However, Saito's notebook was much more efficient. Also, the weights in the cost minimization expression could use the cosine values to reflect points for which one is more confident.\n\n**4 - Improving paths by local search**\nThe output of section 3 is a list of (x, y) positions that define paths.\nI defined another cost function that combined 3 values for each path:\n   - error for (x, y) positions (current positions vs \"starting\" positions, for each point on the path)\n   - error for length of each segment in the path (ratio between length based on start/end positions, and associated step_positions length)\n   - error for heading between consecutive segments (again: variation in heading based on (x, y) positions vs variations in heading based on accelerometer data)\n\nThen, a population-based search was performed using some \"mutation\" operators to update the path and capture constraints to place points in corridors.\nAgain, no CV framework was used here, but improvements in the local search based on this cost function appeared highly correlated with LB scores.\n\n**5 - Post-processing**\nI didn't explore that part that much. I just lazily invoked the snap-to-grid ([https://www.kaggle.com/dragonzhang/3-3-g6-indoor-navigation-snap-to-grid](url)) and leakage ([https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage](url)) notebooks like many of us :-).\nThese notebooks were regularly improving my solutions by a value of around 0.3.\n\n**Final comments**\nOne direction I wanted to investigate was about correcting step_positions returned by the python library provided by the competition host. Indeed, doing some statistics with training data, one could notice significant discrepancies between segments based on actual training positions and corresponding step_positions (roughly, with a factor in the range: 0.5 to 1.5, depending on paths).\n\nI think the main benefit of the approach presented above is that it does not rely too much on training points (except for final snap to grid). I believe this makes it more applicable to real life applications for which positions to predict don't match known positions... However, it does not take advantage enough of the setup of this competition for which a large fraction of test positions matches training positions!\n\n**References:**\nI want to thank the authors of the following 4 notebooks that helped me a lot:\n1) [https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization](url)\nLike many of us, I also used Saito's code in my approach. At the beginning, this notebook was not improving my solution that much, but runtime was much better!\n2) [https://www.kaggle.com/rafaelcartenet/scaled-floors-geojsons-new-dataset](url)\nI wasn't familiar with the shapely library. This notebook provided me with all the code for supporting the local search approach that I implemented. Many thanks to Rafael Cartenet!\n3) [https://www.kaggle.com/dragonzhang/3-3-g6-indoor-navigation-snap-to-grid](url)\n4) [https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage](url)",
      "votes": null
    },
    {
      "id": "1314112",
      "postDate": "05/19/2021 00:49:31",
      "content": "<p>Great solution and congrats on the solo gold!</p>\n<blockquote>\n  <p>no CV…</p>\n</blockquote>\n<p>Love it :D</p>",
      "rawMarkdown": "Great solution and congrats on the solo gold!\n> no CV…\n\nLove it :D",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1314112,
      "author_name": "robikscube",
      "author_url": "",
      "post_date": "05/19/2021 00:49:31",
      "content": "<p>Great solution and congrats on the solo gold!</p>\n<blockquote>\n  <p>no CV…</p>\n</blockquote>\n<p>Love it :D</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1313955": "Congratulations to all the winners. Even if I had no experience in the domain, I found this competition very interesting and decided to enter it very early. This allowed me to try different ideas. Let me present the one that I used in my final submission.\n\n**1 - Pre-processing**\nI used only wifi signals for positions prediction. Beacon or magnetic data didn't help.\nI grouped all wifi rows by block, but I reallocated wifi rows based on \"last seen timestamp\" to the wifi block that was the closest with respect to timestamp.\n\n**2 - Floor prediction and Wifi-based position prediction**\nIn a first step, I used LGBM models for floor prediction (1 model by building) and simple 2 layers feed forward NNs for (x, y) positions (1 model for each floor). The performance was OK.\nThen, I switch to a rather radical approach that was providing better results:\nI computed for each test point the cosine of the angle of wifi fingerprints with respect to each training point. That is: the scal prod divided by the norm of the 2 vectors.\nUsing the cosine instead of the actual scal prod is important because the intensity of the wifi signals varies a lot for different positions.\nAs a result, a perfect match would return a cosine of 1, and the value decreases as the closest training point \"match\" is farther away (actually, there is also a bit of post-processing to discard outlayers).\nFor each test point, the best training point \"match\" is computed FOR EACH FLOOR.\nThe floor prediction is then performed by checking the evolution of the cosine along the path: the floor for which the cosine is the highest more frequently is selected as the predicted floor.\n\nUsing this approach (which is a kind of k-NN), there is no model (no parameter to train), and no CV...\nActually, following the simple idea presented in the thread: [https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization](url), I submitted an edited version of one of my final submissions and was happy to see that my floor prediction was 100% accurate both for public and private data:\n\n| Submission and Description | Private Score | Public Score |\n| --- | --- | --- |\n| submission_df_leak_start_end - REFINED_PATHS-FullPath-SE-NoENF-LS_submission-2021-05-17_07-58-29 - score2.692.csv | 3.25303 | 2.69255 |\n| submission_INCREMENT_2021-05-18_17-42-06.csv | 18.25303 | 17.69255 |\n\n**3 - Introducing accelerometer data**\nWith the approch discussed in section 2, the position prediction is still very noisy. It is necessary to exploit the accelerometer data.\nAt this stage, I didn't start from the raw data but I used instead the library provided by the organisers (\"compute_step_positions\").\nFirst, I used my own post-processing (using some local averaging along paths, combining initial (x, y) predictions and step positions using a moving window). However, Saito's notebook was much more efficient. Also, the weights in the cost minimization expression could use the cosine values to reflect points for which one is more confident.\n\n**4 - Improving paths by local search**\nThe output of section 3 is a list of (x, y) positions that define paths.\nI defined another cost function that combined 3 values for each path:\n   - error for (x, y) positions (current positions vs \"starting\" positions, for each point on the path)\n   - error for length of each segment in the path (ratio between length based on start/end positions, and associated step_positions length)\n   - error for heading between consecutive segments (again: variation in heading based on (x, y) positions vs variations in heading based on accelerometer data)\n\nThen, a population-based search was performed using some \"mutation\" operators to update the path and capture constraints to place points in corridors.\nAgain, no CV framework was used here, but improvements in the local search based on this cost function appeared highly correlated with LB scores.\n\n**5 - Post-processing**\nI didn't explore that part that much. I just lazily invoked the snap-to-grid ([https://www.kaggle.com/dragonzhang/3-3-g6-indoor-navigation-snap-to-grid](url)) and leakage ([https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage](url)) notebooks like many of us :-).\nThese notebooks were regularly improving my solutions by a value of around 0.3.\n\n**Final comments**\nOne direction I wanted to investigate was about correcting step_positions returned by the python library provided by the competition host. Indeed, doing some statistics with training data, one could notice significant discrepancies between segments based on actual training positions and corresponding step_positions (roughly, with a factor in the range: 0.5 to 1.5, depending on paths).\n\nI think the main benefit of the approach presented above is that it does not rely too much on training points (except for final snap to grid). I believe this makes it more applicable to real life applications for which positions to predict don't match known positions... However, it does not take advantage enough of the setup of this competition for which a large fraction of test positions matches training positions!\n\n**References:**\nI want to thank the authors of the following 4 notebooks that helped me a lot:\n1) [https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization](url)\nLike many of us, I also used Saito's code in my approach. At the beginning, this notebook was not improving my solution that much, but runtime was much better!\n2) [https://www.kaggle.com/rafaelcartenet/scaled-floors-geojsons-new-dataset](url)\nI wasn't familiar with the shapely library. This notebook provided me with all the code for supporting the local search approach that I implemented. Many thanks to Rafael Cartenet!\n3) [https://www.kaggle.com/dragonzhang/3-3-g6-indoor-navigation-snap-to-grid](url)\n4) [https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage](url)",
    "1314112": "Great solution and congrats on the solo gold!\n> no CV…\n\nLove it :D"
  },
  "source": "meta"
}