{
  "id": 239882,
  "title": "Quick write up of 5th place",
  "url": "/competitions/indoor-location-navigation/discussion/239882",
  "author_name": "",
  "post_date": "2021-05-18T00:23:02.559647800Z",
  "votes": 48,
  "comment_count": 8,
  "views": 0,
  "content": "<p>This was a very fun competition! I actually have a background in this industry (8 years at an indoor location startup) which is what attracted me to participate. I didn't work on the actual location algorithms there however, so this was mostly all new and interesting.<br>\n<br><br>\nI planned to write a much longer post with more details and diagrams, but I spent all day today trying to improve my score, so I didn't have time to write it today 😆<br>\n<br><br>\n(If you'd like to see more info about any of these steps in a longer post, let me know - I plan to write up more actual details about my models and process)<br>\n<br><br>\nHere are the basics of what I did:<br>\n<br></p>\n<h1>The Process</h1>\n<p><br></p>\n<h2>Step 0. Data exploration</h2>\n<p>I made a mini web app to explore the data, which was incredibly helpful!  Here's a screenshot of what that looked like:<br>\n<img src=\"https://i.imgur.com/1sN94Q1.png\" alt=\"Indoor Dashboard\"><br>\nIs it ugly? yes. 😂 but it's also very functional.<br>\n<br><br>\nIt shows the selected path and all the waypoints for the floor, as well as some sensor data. It also shows the previous end point and next start point of the paths on that floor, as well as the time since the previous end and time until the next path start.<br>\n<br><br>\nAll of that data was incredibly helpful in understanding how the data was collected, and how different sites and floors varied from each other.<br>\n<br></p>\n<h2>Step 1. Predict floors</h2>\n<p>I used a simple CNN based wifi model, combined with some public RNN notebooks (thank you!) and a bit of the device + timestamp leakage I outlined here: <a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/234543\" target=\"_blank\">https://www.kaggle.com/c/indoor-location-navigation/discussion/234543</a> to predict the floor for each path.<br>\n<br><br>\nThe data leakage probably only affected &lt; 1% of the floor predictions, so it was fairly minor I think.\n<br><br>\nI felt pretty good about those floor predictions, so for the remainder of the steps I then kept those predictions constant.<br>\n<br></p>\n<h2>Step 2. Predict approximate path x,y centroid using wifi</h2>\n<p>I tried several RNNs but they either didn't train, or didn't give good results, so I settled on a CNN based model for wifi prediction.<br>\n<br><br>\nI predicted both individual waypoint x,y points as well as the entire path centroid (multiple models) and then ensembled that all with a few of the best public notebooks to get approximate x,y positions for the entire path<br>\n<br></p>\n<h2>Step 3. Predict relative path \"shape\" using sensor data</h2>\n<p>Here's where I think my solution might be a bit unique:<br>\n<br><br>\nI split each path into multiple \"legs\", and then used the sensor data to predict the delta_x and delta_y for each leg individually.  I did this using a CNN for the sensor data + a lot of extra features in a giant model (that I'll detail in a future post), and was able to get a really solid representation of the entire path \"shape\". (down to about +- 1m MAE per leg).<br>\n<br><br>\nI also fine-tuned these models on a per-site basis, which gave it about a 20% boost (according to my training validation split).<br>\n<br><br>\nI was worried that the error would compound when I put it all together (especially for longer paths), but that didn't seem to be the case.<br>\n<br></p>\n<h2>Step 4. Extensive post-processing to fit that path \"shape\" into the floorplan at the approximate wifi x,y location</h2>\n<p>Probably 75% of my time was spent on post-processing, trying to fit the paths (of a predicted shape) onto the floorplan at the predicted wifi x,y point.<br>\n<br><br>\nAgain, I'll write up more details, but some of the post-processing operations I created were:<br>\n<br></p>\n<ul>\n<li>snap to grid</li>\n<li>progressive snap to grid</li>\n<li>slide to hallway</li>\n<li>break up path and slide to hallway</li>\n<li>slide to waypoints</li>\n<li>break up path and slide to waypoints</li>\n<li>simple iterate and snap forwards</li>\n<li>simple iterate and snap reverse</li>\n<li>center between previous stop and next end points</li>\n<li>revert path based on out of bounds relative distance</li>\n<li>revert path based on out of bounds wifi prediction<br>\n<br><br>\n… and more :)<br>\n<br><br>\nI would run various combinations of these post-processing steps, automatically reverting the path if it got too far from the predicted path \"shape\" or predicted wifi x,y point, and then ensembling several runs together to make a submission.<br>\n<br><br>\nOne big boost was when I realized I should split up the longer test paths into multiple \"sub paths\", and average those results.<br>\n<br></li>\n</ul>\n<h2>Step 5. Final submissions</h2>\n<p>My final submission were ensembles of those ensembles - either just averaging, or a few extra rules, including only averaging the closest 2 out of 3 predicted points, or checking the path relative waypoint distance or distance from the predicted wifi.<br>\n<br></p>\n<h1>What went well</h1>\n<p>From my background in indoor location, I knew that uncalibrated wifi would only ever get to about 4m - 5m accuracy, which I think played out in the results.<br>\n<br><br>\nI knew that I would have to switch to sensor data sooner rather than later, and so the decision to do that quickly gave me more time to explore the sensor data.<br>\n<br></p>\n<h1>What I should have done differently</h1>\n<p>I never really got a good cross validation approach working for the post-processing. <br>\n<br><br>\nI was able to do validation for the CNN models I made, but trying to post process the training data never quite worked out.<br>\n<br><br>\nI should have spent more time setting up proper CV for the post-processing of the training data, so I didn't have to rely so much on the public LB for feedback.<br>\n<br></p>\n<h1>Overall</h1>\n<p>I feel like I extracted just about all I could with a dual model + post processing setup… I really tried hard to make a unified single model work, but I kept getting stuck. Good job for any team that made it work!!<br>\n<br></p>\n<h1>Thanks!</h1>\n<p>Thanks to everyone who participated - this was my first Kaggle competition, and a bunch of fun!  I'll post more details in the next few days, and I look forward to competing in the future :)<br>\n<br></p>",
  "messages": [
    {
      "id": "1312262",
      "postDate": "05/18/2021 00:23:02",
      "content": "<p>This was a very fun competition! I actually have a background in this industry (8 years at an indoor location startup) which is what attracted me to participate. I didn't work on the actual location algorithms there however, so this was mostly all new and interesting.<br>\n<br><br>\nI planned to write a much longer post with more details and diagrams, but I spent all day today trying to improve my score, so I didn't have time to write it today 😆<br>\n<br><br>\n(If you'd like to see more info about any of these steps in a longer post, let me know - I plan to write up more actual details about my models and process)<br>\n<br><br>\nHere are the basics of what I did:<br>\n<br></p>\n<h1>The Process</h1>\n<p><br></p>\n<h2>Step 0. Data exploration</h2>\n<p>I made a mini web app to explore the data, which was incredibly helpful!  Here's a screenshot of what that looked like:<br>\n<img src=\"https://i.imgur.com/1sN94Q1.png\" alt=\"Indoor Dashboard\"><br>\nIs it ugly? yes. 😂 but it's also very functional.<br>\n<br><br>\nIt shows the selected path and all the waypoints for the floor, as well as some sensor data. It also shows the previous end point and next start point of the paths on that floor, as well as the time since the previous end and time until the next path start.<br>\n<br><br>\nAll of that data was incredibly helpful in understanding how the data was collected, and how different sites and floors varied from each other.<br>\n<br></p>\n<h2>Step 1. Predict floors</h2>\n<p>I used a simple CNN based wifi model, combined with some public RNN notebooks (thank you!) and a bit of the device + timestamp leakage I outlined here: <a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/234543\" target=\"_blank\">https://www.kaggle.com/c/indoor-location-navigation/discussion/234543</a> to predict the floor for each path.<br>\n<br><br>\nThe data leakage probably only affected &lt; 1% of the floor predictions, so it was fairly minor I think.\n<br><br>\nI felt pretty good about those floor predictions, so for the remainder of the steps I then kept those predictions constant.<br>\n<br></p>\n<h2>Step 2. Predict approximate path x,y centroid using wifi</h2>\n<p>I tried several RNNs but they either didn't train, or didn't give good results, so I settled on a CNN based model for wifi prediction.<br>\n<br><br>\nI predicted both individual waypoint x,y points as well as the entire path centroid (multiple models) and then ensembled that all with a few of the best public notebooks to get approximate x,y positions for the entire path<br>\n<br></p>\n<h2>Step 3. Predict relative path \"shape\" using sensor data</h2>\n<p>Here's where I think my solution might be a bit unique:<br>\n<br><br>\nI split each path into multiple \"legs\", and then used the sensor data to predict the delta_x and delta_y for each leg individually.  I did this using a CNN for the sensor data + a lot of extra features in a giant model (that I'll detail in a future post), and was able to get a really solid representation of the entire path \"shape\". (down to about +- 1m MAE per leg).<br>\n<br><br>\nI also fine-tuned these models on a per-site basis, which gave it about a 20% boost (according to my training validation split).<br>\n<br><br>\nI was worried that the error would compound when I put it all together (especially for longer paths), but that didn't seem to be the case.<br>\n<br></p>\n<h2>Step 4. Extensive post-processing to fit that path \"shape\" into the floorplan at the approximate wifi x,y location</h2>\n<p>Probably 75% of my time was spent on post-processing, trying to fit the paths (of a predicted shape) onto the floorplan at the predicted wifi x,y point.<br>\n<br><br>\nAgain, I'll write up more details, but some of the post-processing operations I created were:<br>\n<br></p>\n<ul>\n<li>snap to grid</li>\n<li>progressive snap to grid</li>\n<li>slide to hallway</li>\n<li>break up path and slide to hallway</li>\n<li>slide to waypoints</li>\n<li>break up path and slide to waypoints</li>\n<li>simple iterate and snap forwards</li>\n<li>simple iterate and snap reverse</li>\n<li>center between previous stop and next end points</li>\n<li>revert path based on out of bounds relative distance</li>\n<li>revert path based on out of bounds wifi prediction<br>\n<br><br>\n… and more :)<br>\n<br><br>\nI would run various combinations of these post-processing steps, automatically reverting the path if it got too far from the predicted path \"shape\" or predicted wifi x,y point, and then ensembling several runs together to make a submission.<br>\n<br><br>\nOne big boost was when I realized I should split up the longer test paths into multiple \"sub paths\", and average those results.<br>\n<br></li>\n</ul>\n<h2>Step 5. Final submissions</h2>\n<p>My final submission were ensembles of those ensembles - either just averaging, or a few extra rules, including only averaging the closest 2 out of 3 predicted points, or checking the path relative waypoint distance or distance from the predicted wifi.<br>\n<br></p>\n<h1>What went well</h1>\n<p>From my background in indoor location, I knew that uncalibrated wifi would only ever get to about 4m - 5m accuracy, which I think played out in the results.<br>\n<br><br>\nI knew that I would have to switch to sensor data sooner rather than later, and so the decision to do that quickly gave me more time to explore the sensor data.<br>\n<br></p>\n<h1>What I should have done differently</h1>\n<p>I never really got a good cross validation approach working for the post-processing. <br>\n<br><br>\nI was able to do validation for the CNN models I made, but trying to post process the training data never quite worked out.<br>\n<br><br>\nI should have spent more time setting up proper CV for the post-processing of the training data, so I didn't have to rely so much on the public LB for feedback.<br>\n<br></p>\n<h1>Overall</h1>\n<p>I feel like I extracted just about all I could with a dual model + post processing setup… I really tried hard to make a unified single model work, but I kept getting stuck. Good job for any team that made it work!!<br>\n<br></p>\n<h1>Thanks!</h1>\n<p>Thanks to everyone who participated - this was my first Kaggle competition, and a bunch of fun!  I'll post more details in the next few days, and I look forward to competing in the future :)<br>\n<br></p>",
      "rawMarkdown": "This was a very fun competition! I actually have a background in this industry (8 years at an indoor location startup) which is what attracted me to participate. I didn't work on the actual location algorithms there however, so this was mostly all new and interesting.\n\n<br />\n\nI planned to write a much longer post with more details and diagrams, but I spent all day today trying to improve my score, so I didn't have time to write it today 😆\n\n<br />\n\n(If you'd like to see more info about any of these steps in a longer post, let me know - I plan to write up more actual details about my models and process)\n\n<br />\n\nHere are the basics of what I did:\n\n<br />\n\n# The Process\n\n<br />\n\n## Step 0. Data exploration\n\nI made a mini web app to explore the data, which was incredibly helpful!  Here's a screenshot of what that looked like:\n\n![Indoor Dashboard](https://i.imgur.com/1sN94Q1.png)\n\nIs it ugly? yes. 😂 but it's also very functional.\n\n<br />\n\nIt shows the selected path and all the waypoints for the floor, as well as some sensor data. It also shows the previous end point and next start point of the paths on that floor, as well as the time since the previous end and time until the next path start.\n\n<br />\n\nAll of that data was incredibly helpful in understanding how the data was collected, and how different sites and floors varied from each other.\n\n<br />\n\n## Step 1. Predict floors\n\nI used a simple CNN based wifi model, combined with some public RNN notebooks (thank you!) and a bit of the device + timestamp leakage I outlined here: https://www.kaggle.com/c/indoor-location-navigation/discussion/234543 to predict the floor for each path.\n<br />\nThe data leakage probably only affected < 1% of the floor predictions, so it was fairly minor I think.\n<br />\nI felt pretty good about those floor predictions, so for the remainder of the steps I then kept those predictions constant.\n\n<br />\n\n## Step 2. Predict approximate path x,y centroid using wifi\n\nI tried several RNNs but they either didn't train, or didn't give good results, so I settled on a CNN based model for wifi prediction.\n<br />\nI predicted both individual waypoint x,y points as well as the entire path centroid (multiple models) and then ensembled that all with a few of the best public notebooks to get approximate x,y positions for the entire path\n\n<br />\n\n## Step 3. Predict relative path \"shape\" using sensor data\n\nHere's where I think my solution might be a bit unique:\n\n<br />\n\nI split each path into multiple \"legs\", and then used the sensor data to predict the delta_x and delta_y for each leg individually.  I did this using a CNN for the sensor data + a lot of extra features in a giant model (that I'll detail in a future post), and was able to get a really solid representation of the entire path \"shape\". (down to about +- 1m MAE per leg).\n\n<br />\n\nI also fine-tuned these models on a per-site basis, which gave it about a 20% boost (according to my training validation split).\n\n<br />\n\nI was worried that the error would compound when I put it all together (especially for longer paths), but that didn't seem to be the case.\n\n<br />\n\n## Step 4. Extensive post-processing to fit that path \"shape\" into the floorplan at the approximate wifi x,y location\n\nProbably 75% of my time was spent on post-processing, trying to fit the paths (of a predicted shape) onto the floorplan at the predicted wifi x,y point.\n<br />\nAgain, I'll write up more details, but some of the post-processing operations I created were:\n\n<br />\n\n- snap to grid\n- progressive snap to grid\n- slide to hallway\n- break up path and slide to hallway\n- slide to waypoints\n- break up path and slide to waypoints\n- simple iterate and snap forwards\n- simple iterate and snap reverse\n- center between previous stop and next end points\n- revert path based on out of bounds relative distance\n- revert path based on out of bounds wifi prediction\n\n<br />\n\n... and more :)\n\n<br />\n\nI would run various combinations of these post-processing steps, automatically reverting the path if it got too far from the predicted path \"shape\" or predicted wifi x,y point, and then ensembling several runs together to make a submission.\n\n<br />\n\nOne big boost was when I realized I should split up the longer test paths into multiple \"sub paths\", and average those results.\n\n<br />\n\n## Step 5. Final submissions\n\nMy final submission were ensembles of those ensembles - either just averaging, or a few extra rules, including only averaging the closest 2 out of 3 predicted points, or checking the path relative waypoint distance or distance from the predicted wifi.\n\n<br />\n\n# What went well\n\nFrom my background in indoor location, I knew that uncalibrated wifi would only ever get to about 4m - 5m accuracy, which I think played out in the results.\n\n<br />\n\nI knew that I would have to switch to sensor data sooner rather than later, and so the decision to do that quickly gave me more time to explore the sensor data.\n\n<br />\n\n\n# What I should have done differently\n\nI never really got a good cross validation approach working for the post-processing. \n\n<br />\n\nI was able to do validation for the CNN models I made, but trying to post process the training data never quite worked out.\n\n<br />\n\nI should have spent more time setting up proper CV for the post-processing of the training data, so I didn't have to rely so much on the public LB for feedback.\n\n<br />\n\n\n# Overall\n\nI feel like I extracted just about all I could with a dual model + post processing setup... I really tried hard to make a unified single model work, but I kept getting stuck. Good job for any team that made it work!!\n\n<br />\n\n# Thanks!\n\nThanks to everyone who participated - this was my first Kaggle competition, and a bunch of fun!  I'll post more details in the next few days, and I look forward to competing in the future :)\n\n<br />",
      "votes": null
    },
    {
      "id": "1312265",
      "postDate": "05/18/2021 00:26:46",
      "content": "<p>thanks for the writeup <a href=\"https://www.kaggle.com/chris62\" target=\"_blank\">@chris62</a> , very interesting. Can you share the score of your model without any postprocessing (or basic) ?</p>",
      "rawMarkdown": "thanks for the writeup @chris62 , very interesting. Can you share the score of your model without any postprocessing (or basic) ?",
      "votes": null
    },
    {
      "id": "1312271",
      "postDate": "05/18/2021 00:30:07",
      "content": "<p>Wifi alone was 4m - 5m (public LB)<br>\nWifi + sensor data was 3m - 4m (public LB)<br>\npost processing brought that down &lt; 3m (public LB)</p>",
      "rawMarkdown": "Wifi alone was 4m - 5m (public LB)\nWifi + sensor data was 3m - 4m (public LB)\npost processing brought that down < 3m (public LB)",
      "votes": null
    },
    {
      "id": "1312361",
      "postDate": "05/18/2021 02:32:00",
      "content": "<blockquote>\n  <p>Is it ugly? yes. 😂 but it's also very functional.</p>\n</blockquote>\n<p>Not ugly at all. It is kicking ass! Congrats!</p>",
      "rawMarkdown": "> Is it ugly? yes. 😂 but it's also very functional.\n\nNot ugly at all. It is kicking ass! Congrats!",
      "votes": null
    },
    {
      "id": "1312383",
      "postDate": "05/18/2021 02:54:22",
      "content": "<p>I totally agree about cv being challenging w/ post processing. A simple example is how every snap-to-grid step had to have separate lists of out of fold waypoints. Similarly, I recalculated leaky waypoints based only on the out-of-fold paths that were adjacent to in fold ones. And then repeat the process for submission paths. Managing this pipeline and tracking all the versions of things took up a major fraction of my time and attention. </p>",
      "rawMarkdown": "I totally agree about cv being challenging w/ post processing. A simple example is how every snap-to-grid step had to have separate lists of out of fold waypoints. Similarly, I recalculated leaky waypoints based only on the out-of-fold paths that were adjacent to in fold ones. And then repeat the process for submission paths. Managing this pipeline and tracking all the versions of things took up a major fraction of my time and attention.",
      "votes": null
    },
    {
      "id": "1312388",
      "postDate": "05/18/2021 02:57:47",
      "content": "<p>I also like that you modeled the shape independently. I suspected shape was very important, but I relied heavily on the “cost min” method and the provided GitHub code as gospel for the shape, I didn’t think to model it myself. Nice work, and congrats!</p>",
      "rawMarkdown": "I also like that you modeled the shape independently. I suspected shape was very important, but I relied heavily on the “cost min” method and the provided GitHub code as gospel for the shape, I didn’t think to model it myself. Nice work, and congrats!",
      "votes": null
    },
    {
      "id": "1312748",
      "postDate": "05/18/2021 08:15:36",
      "content": "<p>This is an interesting write-up <a href=\"https://www.kaggle.com/chris62\" target=\"_blank\">@chris62</a>. Congrats!!</p>",
      "rawMarkdown": "This is an interesting write-up @chris62. Congrats!!",
      "votes": null
    },
    {
      "id": "1312793",
      "postDate": "05/18/2021 08:51:48",
      "content": "<p>Congrats Chris! Looking forward to more details</p>",
      "rawMarkdown": "Congrats Chris! Looking forward to more details",
      "votes": null
    },
    {
      "id": "1314117",
      "postDate": "05/19/2021 00:55:13",
      "content": "<p>Building a webapp to to explore the data 💯. Awesome solution and congrats on the finish.</p>",
      "rawMarkdown": "Building a webapp to to explore the data 💯. Awesome solution and congrats on the finish.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1312265,
      "author_name": "jesucristo",
      "author_url": "",
      "post_date": "05/18/2021 00:26:46",
      "content": "<p>thanks for the writeup <a href=\"https://www.kaggle.com/chris62\" target=\"_blank\">@chris62</a> , very interesting. Can you share the score of your model without any postprocessing (or basic) ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1312271,
          "author_name": "chris62",
          "author_url": "",
          "post_date": "05/18/2021 00:30:07",
          "content": "<p>Wifi alone was 4m - 5m (public LB)<br>\nWifi + sensor data was 3m - 4m (public LB)<br>\npost processing brought that down &lt; 3m (public LB)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1312361,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "05/18/2021 02:32:00",
      "content": "<blockquote>\n  <p>Is it ugly? yes. 😂 but it's also very functional.</p>\n</blockquote>\n<p>Not ugly at all. It is kicking ass! Congrats!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1312383,
      "author_name": "paulfornia",
      "author_url": "",
      "post_date": "05/18/2021 02:54:22",
      "content": "<p>I totally agree about cv being challenging w/ post processing. A simple example is how every snap-to-grid step had to have separate lists of out of fold waypoints. Similarly, I recalculated leaky waypoints based only on the out-of-fold paths that were adjacent to in fold ones. And then repeat the process for submission paths. Managing this pipeline and tracking all the versions of things took up a major fraction of my time and attention. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1312388,
          "author_name": "paulfornia",
          "author_url": "",
          "post_date": "05/18/2021 02:57:47",
          "content": "<p>I also like that you modeled the shape independently. I suspected shape was very important, but I relied heavily on the “cost min” method and the provided GitHub code as gospel for the shape, I didn’t think to model it myself. Nice work, and congrats!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1312748,
      "author_name": "ankitp013",
      "author_url": "",
      "post_date": "05/18/2021 08:15:36",
      "content": "<p>This is an interesting write-up <a href=\"https://www.kaggle.com/chris62\" target=\"_blank\">@chris62</a>. Congrats!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1312793,
      "author_name": "olaf2000",
      "author_url": "",
      "post_date": "05/18/2021 08:51:48",
      "content": "<p>Congrats Chris! Looking forward to more details</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1314117,
      "author_name": "robikscube",
      "author_url": "",
      "post_date": "05/19/2021 00:55:13",
      "content": "<p>Building a webapp to to explore the data 💯. Awesome solution and congrats on the finish.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1312262": "This was a very fun competition! I actually have a background in this industry (8 years at an indoor location startup) which is what attracted me to participate. I didn't work on the actual location algorithms there however, so this was mostly all new and interesting.\n\n<br />\n\nI planned to write a much longer post with more details and diagrams, but I spent all day today trying to improve my score, so I didn't have time to write it today 😆\n\n<br />\n\n(If you'd like to see more info about any of these steps in a longer post, let me know - I plan to write up more actual details about my models and process)\n\n<br />\n\nHere are the basics of what I did:\n\n<br />\n\n# The Process\n\n<br />\n\n## Step 0. Data exploration\n\nI made a mini web app to explore the data, which was incredibly helpful!  Here's a screenshot of what that looked like:\n\n![Indoor Dashboard](https://i.imgur.com/1sN94Q1.png)\n\nIs it ugly? yes. 😂 but it's also very functional.\n\n<br />\n\nIt shows the selected path and all the waypoints for the floor, as well as some sensor data. It also shows the previous end point and next start point of the paths on that floor, as well as the time since the previous end and time until the next path start.\n\n<br />\n\nAll of that data was incredibly helpful in understanding how the data was collected, and how different sites and floors varied from each other.\n\n<br />\n\n## Step 1. Predict floors\n\nI used a simple CNN based wifi model, combined with some public RNN notebooks (thank you!) and a bit of the device + timestamp leakage I outlined here: https://www.kaggle.com/c/indoor-location-navigation/discussion/234543 to predict the floor for each path.\n<br />\nThe data leakage probably only affected < 1% of the floor predictions, so it was fairly minor I think.\n<br />\nI felt pretty good about those floor predictions, so for the remainder of the steps I then kept those predictions constant.\n\n<br />\n\n## Step 2. Predict approximate path x,y centroid using wifi\n\nI tried several RNNs but they either didn't train, or didn't give good results, so I settled on a CNN based model for wifi prediction.\n<br />\nI predicted both individual waypoint x,y points as well as the entire path centroid (multiple models) and then ensembled that all with a few of the best public notebooks to get approximate x,y positions for the entire path\n\n<br />\n\n## Step 3. Predict relative path \"shape\" using sensor data\n\nHere's where I think my solution might be a bit unique:\n\n<br />\n\nI split each path into multiple \"legs\", and then used the sensor data to predict the delta_x and delta_y for each leg individually.  I did this using a CNN for the sensor data + a lot of extra features in a giant model (that I'll detail in a future post), and was able to get a really solid representation of the entire path \"shape\". (down to about +- 1m MAE per leg).\n\n<br />\n\nI also fine-tuned these models on a per-site basis, which gave it about a 20% boost (according to my training validation split).\n\n<br />\n\nI was worried that the error would compound when I put it all together (especially for longer paths), but that didn't seem to be the case.\n\n<br />\n\n## Step 4. Extensive post-processing to fit that path \"shape\" into the floorplan at the approximate wifi x,y location\n\nProbably 75% of my time was spent on post-processing, trying to fit the paths (of a predicted shape) onto the floorplan at the predicted wifi x,y point.\n<br />\nAgain, I'll write up more details, but some of the post-processing operations I created were:\n\n<br />\n\n- snap to grid\n- progressive snap to grid\n- slide to hallway\n- break up path and slide to hallway\n- slide to waypoints\n- break up path and slide to waypoints\n- simple iterate and snap forwards\n- simple iterate and snap reverse\n- center between previous stop and next end points\n- revert path based on out of bounds relative distance\n- revert path based on out of bounds wifi prediction\n\n<br />\n\n... and more :)\n\n<br />\n\nI would run various combinations of these post-processing steps, automatically reverting the path if it got too far from the predicted path \"shape\" or predicted wifi x,y point, and then ensembling several runs together to make a submission.\n\n<br />\n\nOne big boost was when I realized I should split up the longer test paths into multiple \"sub paths\", and average those results.\n\n<br />\n\n## Step 5. Final submissions\n\nMy final submission were ensembles of those ensembles - either just averaging, or a few extra rules, including only averaging the closest 2 out of 3 predicted points, or checking the path relative waypoint distance or distance from the predicted wifi.\n\n<br />\n\n# What went well\n\nFrom my background in indoor location, I knew that uncalibrated wifi would only ever get to about 4m - 5m accuracy, which I think played out in the results.\n\n<br />\n\nI knew that I would have to switch to sensor data sooner rather than later, and so the decision to do that quickly gave me more time to explore the sensor data.\n\n<br />\n\n\n# What I should have done differently\n\nI never really got a good cross validation approach working for the post-processing. \n\n<br />\n\nI was able to do validation for the CNN models I made, but trying to post process the training data never quite worked out.\n\n<br />\n\nI should have spent more time setting up proper CV for the post-processing of the training data, so I didn't have to rely so much on the public LB for feedback.\n\n<br />\n\n\n# Overall\n\nI feel like I extracted just about all I could with a dual model + post processing setup... I really tried hard to make a unified single model work, but I kept getting stuck. Good job for any team that made it work!!\n\n<br />\n\n# Thanks!\n\nThanks to everyone who participated - this was my first Kaggle competition, and a bunch of fun!  I'll post more details in the next few days, and I look forward to competing in the future :)\n\n<br />",
    "1312265": "thanks for the writeup @chris62 , very interesting. Can you share the score of your model without any postprocessing (or basic) ?",
    "1312271": "Wifi alone was 4m - 5m (public LB)\nWifi + sensor data was 3m - 4m (public LB)\npost processing brought that down < 3m (public LB)",
    "1312361": "> Is it ugly? yes. 😂 but it's also very functional.\n\nNot ugly at all. It is kicking ass! Congrats!",
    "1312383": "I totally agree about cv being challenging w/ post processing. A simple example is how every snap-to-grid step had to have separate lists of out of fold waypoints. Similarly, I recalculated leaky waypoints based only on the out-of-fold paths that were adjacent to in fold ones. And then repeat the process for submission paths. Managing this pipeline and tracking all the versions of things took up a major fraction of my time and attention.",
    "1312388": "I also like that you modeled the shape independently. I suspected shape was very important, but I relied heavily on the “cost min” method and the provided GitHub code as gospel for the shape, I didn’t think to model it myself. Nice work, and congrats!",
    "1312748": "This is an interesting write-up @chris62. Congrats!!",
    "1312793": "Congrats Chris! Looking forward to more details",
    "1314117": "Building a webapp to to explore the data 💯. Awesome solution and congrats on the finish."
  },
  "source": "meta"
}