{
  "id": 240151,
  "title": "8th place solution",
  "url": "/competitions/indoor-location-navigation/writeups/micha-stolarczyk-8th-place-solution",
  "author_name": "",
  "post_date": "2021-05-18T17:22:56.290Z",
  "votes": 14,
  "comment_count": 1,
  "views": 0,
  "content": "<h1>Absolute position model</h1>\n<p>My base model used wifi and beacon events to predict absolute position in a single timestamp. It used 50 events from 2 closest blocks of wifi events with hightest values of rssi and 10 closest beacon events. Events were sorted by <code>rssi</code>.</p>\n<p>I trained one model on all sites. I used two bidirectional LSTMs - one on embeddings of wifi bssids + some numerical features related to each wifi event and one on embeddings of beacon_id + some numerical features related to each beacon event. Model also used embeddings of concatenated <code>site</code> and <code>floor</code>.</p>\n<p>I think that the biggest improvement in this model was after I started interpolating waypoints in the training dataset. My best model was trained on waypoints interpolated every 1 second. </p>\n<h1>Relative position model</h1>\n<p>I also trained model, which predicted relative position of a smartphone based on sensor data. It predicted two numbers (deltaX, deltaY) based on 1 second of sensor data - 50 timestamps. I used Kalman Smoother to get targets for this model. It was a simple 1D-CNN model based on raw events.</p>\n<p>Predictions from this model were interpolated, to approximate differences between positions in two consecutive waypoints. It achieved ~2.0 RMSE. </p>\n<h1>Postprocessing</h1>\n<p>The first step of my postprocessing was  inspired by <a href=\"https://www.kaggle.com/saitodevel01\" target=\"_blank\">@saitodevel01</a>'s great notebook: <a href=\"https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization\" target=\"_blank\">indoor - Post-processing by Cost Minimization</a>. </p>\n<p>I assummed that models' predictions have the following distributions</p>\n<p>$$(\\hat{X_i}, \\hat{Y_i}) | (X_i, Y_i) \\sim N_2((X_i, Y_i), \\sigma_1^2 I_2)$$<br>\nand<br>\n$$(\\hat{\\Delta X_i}, \\hat{\\Delta Y_i}) | (X_{i+1}, Y_{i+1}, X_i, Y_i)  \\sim N_2( (X_{i+1}, Y_{i+1}) - (X_i, Y_i), \\sigma_2^2 I_2) $$<br>\nand all pairs \\( (X_i, Y_i) \\) are uniformly distributed in the corridor of the building, where</p>\n<p>\\( (X_i, Y_i) \\)  is a true position at timestamp \\( i \\) , \\( (\\hat{X_i}, \\hat{Y_i}) \\) is a prediction of absolute model, \\(  (\\hat{\\Delta X_i}, \\hat{\\Delta Y_i}) \\) is a prediction of relative model and \\( \\sigma_1, \\sigma_2  \\) are RMSEs of absolute and relative models respectively.</p>\n<p>After some math, it turns out that vector of \\( (X_1, Y_1, …, X_n, Y_n) \\) has truncated multivariate normal distribution with easy to calculate mean and covariance matrix. It is truncated in such a way, that each pair \\( (X_i, Y_i) \\) lies in the corridor of the building.</p>\n<p>I implemented iterative method of sampling paths from this distribution similar to Gibbs sampling. In each iteration it samples pair \\( (X_i, Y_i) \\) using untruncated normal distribution  conditioned on all the other pairs sampled so far until it samples a point inside the corridor or until it reaches given limit of trials. For each path I draw 500 samples and average last 200 samples. </p>\n<p>This method also allowed to use device id leak, by setting \\( \\sigma_1 \\) to some small value for leaked timestamps.</p>\n<p>This method without truncating to corridors had similar performance to Cost Minimiztion notebook  - in fact I think they are almost equivalent. Adding map information improved my public LB score by ~0.3.</p>\n<p>The second step of my postprocessing was standard Snap to Grid, without any additional waypoints.</p>\n<h1>Metrics</h1>\n<table>\n<thead>\n<tr>\n<th>method</th>\n<th>private LB</th>\n<th>public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>best absolute model</td>\n<td>5.36934</td>\n<td>4.90242</td>\n</tr>\n<tr>\n<td>ensemble of 4 absolute models</td>\n<td>5.30067</td>\n<td>4.88062</td>\n</tr>\n<tr>\n<td>above + device id leak + sampling postprocessing</td>\n<td>3.33887</td>\n<td>3.00037</td>\n</tr>\n<tr>\n<td>above + snap to grid</td>\n<td>3.07282</td>\n<td>2.64846</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": "1313722",
      "postDate": "05/18/2021 17:16:31",
      "content": "<h1>Absolute position model</h1>\n<p>My base model used wifi and beacon events to predict absolute position in a single timestamp. It used 50 events from 2 closest blocks of wifi events with hightest values of rssi and 10 closest beacon events. Events were sorted by <code>rssi</code>.</p>\n<p>I trained one model on all sites. I used two bidirectional LSTMs - one on embeddings of wifi bssids + some numerical features related to each wifi event and one on embeddings of beacon_id + some numerical features related to each beacon event. Model also used embeddings of concatenated <code>site</code> and <code>floor</code>.</p>\n<p>I think that the biggest improvement in this model was after I started interpolating waypoints in the training dataset. My best model was trained on waypoints interpolated every 1 second. </p>\n<h1>Relative position model</h1>\n<p>I also trained model, which predicted relative position of a smartphone based on sensor data. It predicted two numbers (deltaX, deltaY) based on 1 second of sensor data - 50 timestamps. I used Kalman Smoother to get targets for this model. It was a simple 1D-CNN model based on raw events.</p>\n<p>Predictions from this model were interpolated, to approximate differences between positions in two consecutive waypoints. It achieved ~2.0 RMSE. </p>\n<h1>Postprocessing</h1>\n<p>The first step of my postprocessing was  inspired by <a href=\"https://www.kaggle.com/saitodevel01\" target=\"_blank\">@saitodevel01</a>'s great notebook: <a href=\"https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization\" target=\"_blank\">indoor - Post-processing by Cost Minimization</a>. </p>\n<p>I assummed that models' predictions have the following distributions</p>\n<p>$$(\\hat{X_i}, \\hat{Y_i}) | (X_i, Y_i) \\sim N_2((X_i, Y_i), \\sigma_1^2 I_2)$$<br>\nand<br>\n$$(\\hat{\\Delta X_i}, \\hat{\\Delta Y_i}) | (X_{i+1}, Y_{i+1}, X_i, Y_i)  \\sim N_2( (X_{i+1}, Y_{i+1}) - (X_i, Y_i), \\sigma_2^2 I_2) $$<br>\nand all pairs \\( (X_i, Y_i) \\) are uniformly distributed in the corridor of the building, where</p>\n<p>\\( (X_i, Y_i) \\)  is a true position at timestamp \\( i \\) , \\( (\\hat{X_i}, \\hat{Y_i}) \\) is a prediction of absolute model, \\(  (\\hat{\\Delta X_i}, \\hat{\\Delta Y_i}) \\) is a prediction of relative model and \\( \\sigma_1, \\sigma_2  \\) are RMSEs of absolute and relative models respectively.</p>\n<p>After some math, it turns out that vector of \\( (X_1, Y_1, …, X_n, Y_n) \\) has truncated multivariate normal distribution with easy to calculate mean and covariance matrix. It is truncated in such a way, that each pair \\( (X_i, Y_i) \\) lies in the corridor of the building.</p>\n<p>I implemented iterative method of sampling paths from this distribution similar to Gibbs sampling. In each iteration it samples pair \\( (X_i, Y_i) \\) using untruncated normal distribution  conditioned on all the other pairs sampled so far until it samples a point inside the corridor or until it reaches given limit of trials. For each path I draw 500 samples and average last 200 samples. </p>\n<p>This method also allowed to use device id leak, by setting \\( \\sigma_1 \\) to some small value for leaked timestamps.</p>\n<p>This method without truncating to corridors had similar performance to Cost Minimiztion notebook  - in fact I think they are almost equivalent. Adding map information improved my public LB score by ~0.3.</p>\n<p>The second step of my postprocessing was standard Snap to Grid, without any additional waypoints.</p>\n<h1>Metrics</h1>\n<table>\n<thead>\n<tr>\n<th>method</th>\n<th>private LB</th>\n<th>public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>best absolute model</td>\n<td>5.36934</td>\n<td>4.90242</td>\n</tr>\n<tr>\n<td>ensemble of 4 absolute models</td>\n<td>5.30067</td>\n<td>4.88062</td>\n</tr>\n<tr>\n<td>above + device id leak + sampling postprocessing</td>\n<td>3.33887</td>\n<td>3.00037</td>\n</tr>\n<tr>\n<td>above + snap to grid</td>\n<td>3.07282</td>\n<td>2.64846</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "# Absolute position model\n\nMy base model used wifi and beacon events to predict absolute position in a single timestamp. It used 50 events from 2 closest blocks of wifi events with hightest values of rssi and 10 closest beacon events. Events were sorted by `rssi`.\n\nI trained one model on all sites. I used two bidirectional LSTMs - one on embeddings of wifi bssids + some numerical features related to each wifi event and one on embeddings of beacon_id + some numerical features related to each beacon event. Model also used embeddings of concatenated `site` and `floor`.\n\nI think that the biggest improvement in this model was after I started interpolating waypoints in the training dataset. My best model was trained on waypoints interpolated every 1 second. \n\n# Relative position model\n\nI also trained model, which predicted relative position of a smartphone based on sensor data. It predicted two numbers (deltaX, deltaY) based on 1 second of sensor data - 50 timestamps. I used Kalman Smoother to get targets for this model. It was a simple 1D-CNN model based on raw events.\n\nPredictions from this model were interpolated, to approximate differences between positions in two consecutive waypoints. It achieved ~2.0 RMSE. \n\n# Postprocessing\n\nThe first step of my postprocessing was  inspired by @saitodevel01's great notebook: [indoor - Post-processing by Cost Minimization](https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization). \n\nI assummed that models' predictions have the following distributions\n\n$$(\\hat{X_i}, \\hat{Y_i}) | (X_i, Y_i) \\sim N_2((X_i, Y_i), \\sigma_1^2 I_2)$$\nand\n$$(\\hat{\\Delta X_i}, \\hat{\\Delta Y_i}) | (X_{i+1}, Y_{i+1}, X_i, Y_i)  \\sim N_2( (X_{i+1}, Y_{i+1}) - (X_i, Y_i), \\sigma_2^2 I_2) $$\nand all pairs \\\\( (X_i, Y_i) \\\\) are uniformly distributed in the corridor of the building, where\n\n\\\\( (X_i, Y_i) \\\\)  is a true position at timestamp \\\\( i \\\\) , \\\\( (\\hat{X_i}, \\hat{Y_i}) \\\\) is a prediction of absolute model, \\\\(  (\\hat{\\Delta X_i}, \\hat{\\Delta Y_i}) \\\\) is a prediction of relative model and \\\\( \\sigma_1, \\sigma_2  \\\\) are RMSEs of absolute and relative models respectively.\n\nAfter some math, it turns out that vector of \\\\( (X_1, Y_1, ..., X_n, Y_n) \\\\) has truncated multivariate normal distribution with easy to calculate mean and covariance matrix. It is truncated in such a way, that each pair \\\\( (X_i, Y_i) \\\\) lies in the corridor of the building.\n\nI implemented iterative method of sampling paths from this distribution similar to Gibbs sampling. In each iteration it samples pair \\\\( (X_i, Y_i) \\\\) using untruncated normal distribution  conditioned on all the other pairs sampled so far until it samples a point inside the corridor or until it reaches given limit of trials. For each path I draw 500 samples and average last 200 samples. \n\nThis method also allowed to use device id leak, by setting \\\\( \\sigma_1 \\\\) to some small value for leaked timestamps.\n\nThis method without truncating to corridors had similar performance to Cost Minimiztion notebook  - in fact I think they are almost equivalent. Adding map information improved my public LB score by ~0.3.\n\nThe second step of my postprocessing was standard Snap to Grid, without any additional waypoints.\n\n# Metrics\n\n| method | private LB | public LB |\n| --- | --- | --- |\n| best absolute model |5.36934 |4.90242 |\n| ensemble of 4 absolute models |5.30067 | 4.88062|\n| above + device id leak + sampling postprocessing |3.33887  |  3.00037|\n| above + snap to grid |3.07282 | 2.64846 |",
      "votes": null
    },
    {
      "id": "1314113",
      "postDate": "05/19/2021 00:51:25",
      "content": "<p>Congrats! Great solution and solo gold.</p>",
      "rawMarkdown": "Congrats! Great solution and solo gold.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1314113,
      "author_name": "robikscube",
      "author_url": "",
      "post_date": "05/19/2021 00:51:25",
      "content": "<p>Congrats! Great solution and solo gold.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1313722": "# Absolute position model\n\nMy base model used wifi and beacon events to predict absolute position in a single timestamp. It used 50 events from 2 closest blocks of wifi events with hightest values of rssi and 10 closest beacon events. Events were sorted by `rssi`.\n\nI trained one model on all sites. I used two bidirectional LSTMs - one on embeddings of wifi bssids + some numerical features related to each wifi event and one on embeddings of beacon_id + some numerical features related to each beacon event. Model also used embeddings of concatenated `site` and `floor`.\n\nI think that the biggest improvement in this model was after I started interpolating waypoints in the training dataset. My best model was trained on waypoints interpolated every 1 second. \n\n# Relative position model\n\nI also trained model, which predicted relative position of a smartphone based on sensor data. It predicted two numbers (deltaX, deltaY) based on 1 second of sensor data - 50 timestamps. I used Kalman Smoother to get targets for this model. It was a simple 1D-CNN model based on raw events.\n\nPredictions from this model were interpolated, to approximate differences between positions in two consecutive waypoints. It achieved ~2.0 RMSE. \n\n# Postprocessing\n\nThe first step of my postprocessing was  inspired by @saitodevel01's great notebook: [indoor - Post-processing by Cost Minimization](https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization). \n\nI assummed that models' predictions have the following distributions\n\n$$(\\hat{X_i}, \\hat{Y_i}) | (X_i, Y_i) \\sim N_2((X_i, Y_i), \\sigma_1^2 I_2)$$\nand\n$$(\\hat{\\Delta X_i}, \\hat{\\Delta Y_i}) | (X_{i+1}, Y_{i+1}, X_i, Y_i)  \\sim N_2( (X_{i+1}, Y_{i+1}) - (X_i, Y_i), \\sigma_2^2 I_2) $$\nand all pairs \\\\( (X_i, Y_i) \\\\) are uniformly distributed in the corridor of the building, where\n\n\\\\( (X_i, Y_i) \\\\)  is a true position at timestamp \\\\( i \\\\) , \\\\( (\\hat{X_i}, \\hat{Y_i}) \\\\) is a prediction of absolute model, \\\\(  (\\hat{\\Delta X_i}, \\hat{\\Delta Y_i}) \\\\) is a prediction of relative model and \\\\( \\sigma_1, \\sigma_2  \\\\) are RMSEs of absolute and relative models respectively.\n\nAfter some math, it turns out that vector of \\\\( (X_1, Y_1, ..., X_n, Y_n) \\\\) has truncated multivariate normal distribution with easy to calculate mean and covariance matrix. It is truncated in such a way, that each pair \\\\( (X_i, Y_i) \\\\) lies in the corridor of the building.\n\nI implemented iterative method of sampling paths from this distribution similar to Gibbs sampling. In each iteration it samples pair \\\\( (X_i, Y_i) \\\\) using untruncated normal distribution  conditioned on all the other pairs sampled so far until it samples a point inside the corridor or until it reaches given limit of trials. For each path I draw 500 samples and average last 200 samples. \n\nThis method also allowed to use device id leak, by setting \\\\( \\sigma_1 \\\\) to some small value for leaked timestamps.\n\nThis method without truncating to corridors had similar performance to Cost Minimiztion notebook  - in fact I think they are almost equivalent. Adding map information improved my public LB score by ~0.3.\n\nThe second step of my postprocessing was standard Snap to Grid, without any additional waypoints.\n\n# Metrics\n\n| method | private LB | public LB |\n| --- | --- | --- |\n| best absolute model |5.36934 |4.90242 |\n| ensemble of 4 absolute models |5.30067 | 4.88062|\n| above + device id leak + sampling postprocessing |3.33887  |  3.00037|\n| above + snap to grid |3.07282 | 2.64846 |",
    "1314113": "Congrats! Great solution and solo gold."
  },
  "source": "meta"
}