{
  "id": 234543,
  "title": "Known data leaks so far (leveling the field)",
  "url": "/competitions/indoor-location-navigation/discussion/234543",
  "author_name": "",
  "post_date": "2021-04-24T21:13:08.342928400Z",
  "votes": 78,
  "comment_count": 29,
  "views": 0,
  "content": "<p>There have been previous posts about some of these data leaks, but in order to level the playing field, I wanted to outline the 3 data leaks I know about, and how I'm using them.</p>\n<p>I don't share this information lightly, but I think it's important that the competition be fair, and I don't want unintentional data leaks to be the reason for victory :)</p>\n<p>I am sorry if some of the top teams are using these leaks secretly - but hopefully you understand why I'm publishing them.</p>\n<p><strong>1. Timestamp leakage</strong></p>\n<p>The event timestamps have been reset in the test data files, but the <em>wifi last seen timetstamps have not</em></p>\n<p>If you look at a TYPE_WIFI record from a test path:</p>\n<p><code>0000000001900 TYPE_WIFI 09f836894fc1fe9af6f429fc24dcccc2e6847fe0  b773f508054e9d212d9ef69fea73222c20eece4c  -52 2462  1573797536775</code></p>\n<p>The last value there is the true timestamp of the \"last seen\" wifi entry, even though the first value is relative to the start of the path.</p>\n<p>This has been already reported:</p>\n<p><a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/228898\" target=\"_blank\">https://www.kaggle.com/c/indoor-location-navigation/discussion/228898</a></p>\n<p><a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/232072\" target=\"_blank\">https://www.kaggle.com/c/indoor-location-navigation/discussion/232072</a></p>\n<p>That means that <em>you can slot all the test paths into the time where they were collected</em> relative to the training paths.</p>\n<p><em>How I'm using it:</em> I was very excited about this at first - but honestly, it isn't as big as it seems at first, since you can get about 99% of the way there with a good model anyway. However, the biggest place this can be used is for <em>floor level calculations</em>.</p>\n<p>The test paths are taken from a sampling of all gathered paths, and often, those test paths are taken right from the middle of a large group of training paths that are all on the same floor.</p>\n<p>Now - a simple wifi model should be able to capture 99% of the floors correctly anyway.</p>\n<p>Here's one specific example from my test set though:</p>\n<p>For one test path, my model was predicting F4, but based on the timestamp when it occurred, it was very clear that the path <em>must</em> have come from F3.</p>\n<p>It is fairly simple (once you see some of these paths) to create a few simple rules about what floor it must have come from, based on which floor the surrounding paths are on.</p>\n<p><em>A word of caution:</em> don't assume that all test paths are located on the training floors that surround them, and don't assume that only one device was collecting information at any one time!</p>\n<p><strong>2. End point is the same as some start points</strong></p>\n<p>This has already been reported in a few posts:</p>\n<p><a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/228898\" target=\"_blank\">https://www.kaggle.com/c/indoor-location-navigation/discussion/228898</a></p>\n<p><a href=\"https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage\" target=\"_blank\">https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage</a></p>\n<p>The idea is: some paths have the exact same stop/end waypoints as the previous or next paths.</p>\n<p>If you combine that with my leak #1, that means that for some posts, you can know exactly what the start or stop points are.</p>\n<p>A few notes about this data leak: A) it only applies to some of the test paths, but B) if you generally use that idea combined with leak #1, that test paths close in time to training paths will generally be in similar areas of a site/floor - then you can sometimes correlate the end or start of training paths with the start or end of test paths…</p>\n<p>Some of that should be obvious: if one person is recording both training and test paths, then if they only have a few seconds between paths, they could have only traveled a few meters…</p>\n<p><strong>How I'm using it</strong>: this data leak is what really caused me to write this post; noticing and incorporating this is what dropped me from 3.x+ to 2.x. I didn't feel good about how well that worked, so I'm writing this post to let others know about the data leak.</p>\n<p><strong>3. Device leakage</strong></p>\n<p>This is the data leak that I don't think has been talked about on this forum before.</p>\n<p>If you look at the uncalibrated sensor information, you might see a row like:</p>\n<p><code>0000000000276 TYPE_MAGNETIC_FIELD_UNCALIBRATED  -106.762695 35.142517 -355.44434  -83.41217 16.18042  -325.82092  3</code></p>\n<p>Those last 3 float values (before the accuracy) are <strong>device specific</strong>. Since the same devices are used for the training and test paths, it SEEMS like this is a huge leak - because combining with leak #1, I can almost exactly specify what device was recording which path at what time. (and so exactly specify the test floors)</p>\n<p>Again, this seems like a huge leakage, but when combined with #1 and #2, is actually secondary… but can still be very helpful as a meta feature to certain models.</p>\n<p><strong>How I'm using this</strong>: I'm using this \"device\" feature in my models, so I'm a bit unsure about how little or much this is influencing my results. If you are using the uncalibrated sensor values, then you should see about as much improvement as I'm seeing anyway though.</p>\n<p>I know this can help a bit with floor level predictions - but again, if you have a generally good wifi/beacon model, then this data leak probably doesn't matter that much.</p>\n<h1>Why post this</h1>\n<p>Hopefully this discussion post has provided some insight about the competition… again - I'm sorry if some of the top teams were using these leaks in secret, but I think that it is a much more fair competition if the leaks are made public.</p>\n<p>In my own models, these leaks have undoubtedly provided some boost - but I'd really like to see the winning models without the leak advantages - so hopefully this helps all teams.</p>\n<p>Thanks,<br>\nChris</p>",
  "messages": [
    {
      "id": "1283355",
      "postDate": "04/24/2021 21:13:08",
      "content": "<p>There have been previous posts about some of these data leaks, but in order to level the playing field, I wanted to outline the 3 data leaks I know about, and how I'm using them.</p>\n<p>I don't share this information lightly, but I think it's important that the competition be fair, and I don't want unintentional data leaks to be the reason for victory :)</p>\n<p>I am sorry if some of the top teams are using these leaks secretly - but hopefully you understand why I'm publishing them.</p>\n<p><strong>1. Timestamp leakage</strong></p>\n<p>The event timestamps have been reset in the test data files, but the <em>wifi last seen timetstamps have not</em></p>\n<p>If you look at a TYPE_WIFI record from a test path:</p>\n<p><code>0000000001900 TYPE_WIFI 09f836894fc1fe9af6f429fc24dcccc2e6847fe0  b773f508054e9d212d9ef69fea73222c20eece4c  -52 2462  1573797536775</code></p>\n<p>The last value there is the true timestamp of the \"last seen\" wifi entry, even though the first value is relative to the start of the path.</p>\n<p>This has been already reported:</p>\n<p><a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/228898\" target=\"_blank\">https://www.kaggle.com/c/indoor-location-navigation/discussion/228898</a></p>\n<p><a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/232072\" target=\"_blank\">https://www.kaggle.com/c/indoor-location-navigation/discussion/232072</a></p>\n<p>That means that <em>you can slot all the test paths into the time where they were collected</em> relative to the training paths.</p>\n<p><em>How I'm using it:</em> I was very excited about this at first - but honestly, it isn't as big as it seems at first, since you can get about 99% of the way there with a good model anyway. However, the biggest place this can be used is for <em>floor level calculations</em>.</p>\n<p>The test paths are taken from a sampling of all gathered paths, and often, those test paths are taken right from the middle of a large group of training paths that are all on the same floor.</p>\n<p>Now - a simple wifi model should be able to capture 99% of the floors correctly anyway.</p>\n<p>Here's one specific example from my test set though:</p>\n<p>For one test path, my model was predicting F4, but based on the timestamp when it occurred, it was very clear that the path <em>must</em> have come from F3.</p>\n<p>It is fairly simple (once you see some of these paths) to create a few simple rules about what floor it must have come from, based on which floor the surrounding paths are on.</p>\n<p><em>A word of caution:</em> don't assume that all test paths are located on the training floors that surround them, and don't assume that only one device was collecting information at any one time!</p>\n<p><strong>2. End point is the same as some start points</strong></p>\n<p>This has already been reported in a few posts:</p>\n<p><a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/228898\" target=\"_blank\">https://www.kaggle.com/c/indoor-location-navigation/discussion/228898</a></p>\n<p><a href=\"https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage\" target=\"_blank\">https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage</a></p>\n<p>The idea is: some paths have the exact same stop/end waypoints as the previous or next paths.</p>\n<p>If you combine that with my leak #1, that means that for some posts, you can know exactly what the start or stop points are.</p>\n<p>A few notes about this data leak: A) it only applies to some of the test paths, but B) if you generally use that idea combined with leak #1, that test paths close in time to training paths will generally be in similar areas of a site/floor - then you can sometimes correlate the end or start of training paths with the start or end of test paths…</p>\n<p>Some of that should be obvious: if one person is recording both training and test paths, then if they only have a few seconds between paths, they could have only traveled a few meters…</p>\n<p><strong>How I'm using it</strong>: this data leak is what really caused me to write this post; noticing and incorporating this is what dropped me from 3.x+ to 2.x. I didn't feel good about how well that worked, so I'm writing this post to let others know about the data leak.</p>\n<p><strong>3. Device leakage</strong></p>\n<p>This is the data leak that I don't think has been talked about on this forum before.</p>\n<p>If you look at the uncalibrated sensor information, you might see a row like:</p>\n<p><code>0000000000276 TYPE_MAGNETIC_FIELD_UNCALIBRATED  -106.762695 35.142517 -355.44434  -83.41217 16.18042  -325.82092  3</code></p>\n<p>Those last 3 float values (before the accuracy) are <strong>device specific</strong>. Since the same devices are used for the training and test paths, it SEEMS like this is a huge leak - because combining with leak #1, I can almost exactly specify what device was recording which path at what time. (and so exactly specify the test floors)</p>\n<p>Again, this seems like a huge leakage, but when combined with #1 and #2, is actually secondary… but can still be very helpful as a meta feature to certain models.</p>\n<p><strong>How I'm using this</strong>: I'm using this \"device\" feature in my models, so I'm a bit unsure about how little or much this is influencing my results. If you are using the uncalibrated sensor values, then you should see about as much improvement as I'm seeing anyway though.</p>\n<p>I know this can help a bit with floor level predictions - but again, if you have a generally good wifi/beacon model, then this data leak probably doesn't matter that much.</p>\n<h1>Why post this</h1>\n<p>Hopefully this discussion post has provided some insight about the competition… again - I'm sorry if some of the top teams were using these leaks in secret, but I think that it is a much more fair competition if the leaks are made public.</p>\n<p>In my own models, these leaks have undoubtedly provided some boost - but I'd really like to see the winning models without the leak advantages - so hopefully this helps all teams.</p>\n<p>Thanks,<br>\nChris</p>",
      "rawMarkdown": "There have been previous posts about some of these data leaks, but in order to level the playing field, I wanted to outline the 3 data leaks I know about, and how I'm using them.\n\nI don't share this information lightly, but I think it's important that the competition be fair, and I don't want unintentional data leaks to be the reason for victory :)\n\nI am sorry if some of the top teams are using these leaks secretly - but hopefully you understand why I'm publishing them.\n\n\n**1. Timestamp leakage**\n\nThe event timestamps have been reset in the test data files, but the _wifi last seen timetstamps have not_\n\nIf you look at a TYPE_WIFI record from a test path:\n\n`0000000001900 TYPE_WIFI 09f836894fc1fe9af6f429fc24dcccc2e6847fe0  b773f508054e9d212d9ef69fea73222c20eece4c  -52 2462  1573797536775`\n\nThe last value there is the true timestamp of the \"last seen\" wifi entry, even though the first value is relative to the start of the path.\n\nThis has been already reported:\n\nhttps://www.kaggle.com/c/indoor-location-navigation/discussion/228898\n\nhttps://www.kaggle.com/c/indoor-location-navigation/discussion/232072\n\nThat means that *you can slot all the test paths into the time where they were collected* relative to the training paths.\n\n*How I'm using it:* I was very excited about this at first - but honestly, it isn't as big as it seems at first, since you can get about 99% of the way there with a good model anyway. However, the biggest place this can be used is for *floor level calculations*.\n\nThe test paths are taken from a sampling of all gathered paths, and often, those test paths are taken right from the middle of a large group of training paths that are all on the same floor.\n\nNow - a simple wifi model should be able to capture 99% of the floors correctly anyway.\n\nHere's one specific example from my test set though:\n\nFor one test path, my model was predicting F4, but based on the timestamp when it occurred, it was very clear that the path _must_ have come from F3.\n\nIt is fairly simple (once you see some of these paths) to create a few simple rules about what floor it must have come from, based on which floor the surrounding paths are on.\n\n*A word of caution:* don't assume that all test paths are located on the training floors that surround them, and don't assume that only one device was collecting information at any one time!\n\n\n**2. End point is the same as some start points**\n\nThis has already been reported in a few posts:\n\nhttps://www.kaggle.com/c/indoor-location-navigation/discussion/228898\n\nhttps://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage\n\nThe idea is: some paths have the exact same stop/end waypoints as the previous or next paths.\n\nIf you combine that with my leak #1, that means that for some posts, you can know exactly what the start or stop points are.\n\nA few notes about this data leak: A) it only applies to some of the test paths, but B) if you generally use that idea combined with leak #1, that test paths close in time to training paths will generally be in similar areas of a site/floor - then you can sometimes correlate the end or start of training paths with the start or end of test paths...\n\nSome of that should be obvious: if one person is recording both training and test paths, then if they only have a few seconds between paths, they could have only traveled a few meters...\n\n**How I'm using it**: this data leak is what really caused me to write this post; noticing and incorporating this is what dropped me from 3.x+ to 2.x. I didn't feel good about how well that worked, so I'm writing this post to let others know about the data leak.\n\n\n**3. Device leakage**\n\nThis is the data leak that I don't think has been talked about on this forum before.\n\nIf you look at the uncalibrated sensor information, you might see a row like:\n\n`0000000000276 TYPE_MAGNETIC_FIELD_UNCALIBRATED  -106.762695 35.142517 -355.44434  -83.41217 16.18042  -325.82092  3`\n\nThose last 3 float values (before the accuracy) are **device specific**. Since the same devices are used for the training and test paths, it SEEMS like this is a huge leak - because combining with leak #1, I can almost exactly specify what device was recording which path at what time. (and so exactly specify the test floors)\n\nAgain, this seems like a huge leakage, but when combined with #1 and #2, is actually secondary... but can still be very helpful as a meta feature to certain models.\n\n**How I'm using this**: I'm using this \"device\" feature in my models, so I'm a bit unsure about how little or much this is influencing my results. If you are using the uncalibrated sensor values, then you should see about as much improvement as I'm seeing anyway though.\n\nI know this can help a bit with floor level predictions - but again, if you have a generally good wifi/beacon model, then this data leak probably doesn't matter that much.\n\n\n# Why post this\n\nHopefully this discussion post has provided some insight about the competition... again - I'm sorry if some of the top teams were using these leaks in secret, but I think that it is a much more fair competition if the leaks are made public.\n\nIn my own models, these leaks have undoubtedly provided some boost - but I'd really like to see the winning models without the leak advantages - so hopefully this helps all teams.\n\n\n\nThanks,\nChris",
      "votes": null
    },
    {
      "id": "1283397",
      "postDate": "04/24/2021 22:17:18",
      "content": "<p>Thanks for sharing and being generous and fair!<br>\nI'll try to utilize the device leakage to see if it works for us.</p>",
      "rawMarkdown": "Thanks for sharing and being generous and fair!\nI'll try to utilize the device leakage to see if it works for us.",
      "votes": null
    },
    {
      "id": "1283438",
      "postDate": "04/24/2021 23:36:45",
      "content": "<p>I learned something new from this discussion.I would like to reflect them in my prediction.<br>\nThank you for sharing leaks with us.</p>",
      "rawMarkdown": "I learned something new from this discussion.I would like to reflect them in my prediction.\nThank you for sharing leaks with us.",
      "votes": null
    },
    {
      "id": "1283442",
      "postDate": "04/24/2021 23:40:22",
      "content": "<p>Thank you for sharing them. I did not notice 3rd point.<br>\nWould you elaborate why you think the last 3 float values in uncalibrated sensor information are device specific?<br>\nAre they consistent with waypoint leakage (2nd point) and/or <a href=\"https://www.kaggle.com/tomooinubushi/retrieving-user-id-from-leaked-wifi-feature\" target=\"_blank\">shared wifi records leakage</a>?</p>",
      "rawMarkdown": "Thank you for sharing them. I did not notice 3rd point.\nWould you elaborate why you think the last 3 float values in uncalibrated sensor information are device specific?\nAre they consistent with waypoint leakage (2nd point) and/or [shared wifi records leakage](https://www.kaggle.com/tomooinubushi/retrieving-user-id-from-leaked-wifi-feature)?",
      "votes": null
    },
    {
      "id": "1283447",
      "postDate": "04/24/2021 23:47:07",
      "content": "<p>The uncalibrated sensors are supplied through Android's \"uncalibrated sensors\" api: <a href=\"https://source.android.com/devices/sensors/sensor-types#uncalibrated_sensors\" target=\"_blank\">https://source.android.com/devices/sensors/sensor-types#uncalibrated_sensors</a></p>\n<p>The \"bias\" terms (the second set of 3 values) are set through the device calibration; so they are the same (or similar) for each device (though they drift over time).</p>\n<p>That doesn't mean they're identical for each device - but if you look at two different devices, they will probably be quite different; so if they are similar across two timestamps, it can be assumed to be the same device.</p>\n<p>Hope that helps!</p>",
      "rawMarkdown": "The uncalibrated sensors are supplied through Android's \"uncalibrated sensors\" api: https://source.android.com/devices/sensors/sensor-types#uncalibrated_sensors\n\nThe \"bias\" terms (the second set of 3 values) are set through the device calibration; so they are the same (or similar) for each device (though they drift over time).\n\nThat doesn't mean they're identical for each device - but if you look at two different devices, they will probably be quite different; so if they are similar across two timestamps, it can be assumed to be the same device.\n\nHope that helps!",
      "votes": null
    },
    {
      "id": "1283452",
      "postDate": "04/24/2021 23:56:41",
      "content": "<p>Thank you for your reply. I agree with your conclusion. <br>\nI hope this is the last leakage we can find in this competition.</p>",
      "rawMarkdown": "Thank you for your reply. I agree with your conclusion. \nI hope this is the last leakage we can find in this competition.",
      "votes": null
    },
    {
      "id": "1283470",
      "postDate": "04/25/2021 00:47:05",
      "content": "<p>Thank you so much for posting this. Regarding this point:</p>\n<blockquote>\n  <p>How I'm using it: this data leak is what really caused me to write this post; noticing and incorporating this is what dropped me from 3.x+ to 2.x. I didn't feel good about how well that worked, so I'm writing this post to let others know about the data leak.</p>\n</blockquote>\n<p>I remember the <a href=\"https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage\" target=\"_blank\">kernel</a>, which uses the similar idea of leak 1)+2), mentioned that the improvement was 0.1. Do you mean it can actually be 1.0? Thank you!</p>",
      "rawMarkdown": "Thank you so much for posting this. Regarding this point:\n>How I'm using it: this data leak is what really caused me to write this post; noticing and incorporating this is what dropped me from 3.x+ to 2.x. I didn't feel good about how well that worked, so I'm writing this post to let others know about the data leak.\n\nI remember the [kernel](https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage), which uses the similar idea of leak 1)+2), mentioned that the improvement was 0.1. Do you mean it can actually be 1.0? Thank you!",
      "votes": null
    },
    {
      "id": "1283475",
      "postDate": "04/25/2021 00:51:46",
      "content": "<p>If you just assign all the start waypoints to the previous end waypoint, and all the end waypoints to the next start waypoints, then you'll do as much harm as good, and won't get much boost :)</p>\n<p>But if you are more careful with when and how you use this info, then you can do better than 0.1</p>\n<p>It's difficult to say in my case because I made several improvements all at once - but yes, this is one of the factors I credit with my results going from 3.5 to 2.5. Perhaps it was half of that difference? (so about 0.5?) - difficult to say.</p>",
      "rawMarkdown": "If you just assign all the start waypoints to the previous end waypoint, and all the end waypoints to the next start waypoints, then you'll do as much harm as good, and won't get much boost :)\n\nBut if you are more careful with when and how you use this info, then you can do better than 0.1\n\nIt's difficult to say in my case because I made several improvements all at once - but yes, this is one of the factors I credit with my results going from 3.5 to 2.5. Perhaps it was half of that difference? (so about 0.5?) - difficult to say.",
      "votes": null
    },
    {
      "id": "1283493",
      "postDate": "04/25/2021 01:32:27",
      "content": "<p><a href=\"https://www.kaggle.com/chris62\" target=\"_blank\">@chris62</a> <br>\nJust FYI and I guess you are aware of it but there seems multiple pseudo device ids in one path file.</p>\n<pre><code>% grep TYPE_MAGNETIC_FIELD_UNCALIBRATED  5dcf888d878f3300066c6e42.txt | cut -f6-8 | uniq\n-2.3880005      1.7684937       -316.80603\n-5.015564       3.1524658       -391.75568\n</code></pre>",
      "rawMarkdown": "chris62 \nJust FYI and I guess you are aware of it but there seems multiple pseudo device ids in one path file.\n\n```\n% grep TYPE_MAGNETIC_FIELD_UNCALIBRATED  5dcf888d878f3300066c6e42.txt | cut -f6-8 | uniq\n-2.3880005      1.7684937       -316.80603\n-5.015564       3.1524658       -391.75568\n```",
      "votes": null
    },
    {
      "id": "1283500",
      "postDate": "04/25/2021 01:40:34",
      "content": "<p>Yes, if the device is recalibrated mid-path, then the bias terms will change.  I don't remember seeing a drift that large, but guess I missed it :), but yes, it will \"drift\" or change over time for the same device, so I had to accept some % difference as still the same device (though totally separate devices are generally easy to separate, which is where I think it's probably most useful)</p>",
      "rawMarkdown": "Yes, if the device is recalibrated mid-path, then the bias terms will change.  I don't remember seeing a drift that large, but guess I missed it :), but yes, it will \"drift\" or change over time for the same device, so I had to accept some % difference as still the same device (though totally separate devices are generally easy to separate, which is where I think it's probably most useful)",
      "votes": null
    },
    {
      "id": "1283503",
      "postDate": "04/25/2021 01:44:20",
      "content": "<p>Thanks. The \"drift\" is not happening that often so I think it's should be fine :) </p>",
      "rawMarkdown": "Thanks. The \"drift\" is not happening that often so I think it's should be fine :)",
      "votes": null
    },
    {
      "id": "1283516",
      "postDate": "04/25/2021 02:19:33",
      "content": "<p>I agree with <a href=\"https://www.kaggle.com/chris62\" target=\"_blank\">@chris62</a>. I should also add that my postprocessing only processed 176 start/end waypoints of 10,133 test waypoints (1.7%) and got 0.07 gain.<br>\nIf you use these information as input of path-wise time series models like lstm, for example, the leaked information would propagate to the predictions of second and third waypoints and might cause broader changes. <br>\nOne more thing I should add is I did not include device/user wise information to estimate leaked start/end waypoints in that kernel. I thought nobody use this postprocessing as is, the fact behind it is important. </p>",
      "rawMarkdown": "I agree with @chris62. I should also add that my postprocessing only processed 176 start/end waypoints of 10,133 test waypoints (1.7%) and got 0.07 gain.\nIf you use these information as input of path-wise time series models like lstm, for example, the leaked information would propagate to the predictions of second and third waypoints and might cause broader changes. \nOne more thing I should add is I did not include device/user wise information to estimate leaked start/end waypoints in that kernel. I thought nobody use this postprocessing as is, the fact behind it is important.",
      "votes": null
    },
    {
      "id": "1283564",
      "postDate": "04/25/2021 03:55:56",
      "content": "<p>Thank you for sharing. As far as I know, Our team doesn't use raw timestamp information (#1) and device information (#3) now. Youri tried to use the start/end point leakage (#2) by <a href=\"https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage\" target=\"_blank\">tomoo's kernel</a> 4 days ago, but he said the model improved from 2.715 to 2.712 (only 0.003 improvement). <br>\nSo, We don't use any leakage explicitly now. However, as <a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/231311\" target=\"_blank\">tomoo said</a>, probably we all already used this leakage by wifi. This is probably because the improvement is small.</p>",
      "rawMarkdown": "Thank you for sharing. As far as I know, Our team doesn't use raw timestamp information (#1) and device information (#3) now. Youri tried to use the start/end point leakage (#2) by [tomoo's kernel](https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage) 4 days ago, but he said the model improved from 2.715 to 2.712 (only 0.003 improvement). \nSo, We don't use any leakage explicitly now. However, as [tomoo said](https://www.kaggle.com/c/indoor-location-navigation/discussion/231311), probably we all already used this leakage by wifi. This is probably because the improvement is small.",
      "votes": null
    },
    {
      "id": "1283737",
      "postDate": "04/25/2021 08:03:02",
      "content": "<p>Thank you both. Very helpful.</p>",
      "rawMarkdown": "Thank you both. Very helpful.",
      "votes": null
    },
    {
      "id": "1284411",
      "postDate": "04/25/2021 22:08:30",
      "content": "<p>I found 5e034414ae83d40006039063 has 158 pseudo device ids 👀</p>",
      "rawMarkdown": "I found 5e034414ae83d40006039063 has 158 pseudo device ids 👀",
      "votes": null
    },
    {
      "id": "1284414",
      "postDate": "04/25/2021 22:14:02",
      "content": "<p>How different is each one? The android sensors seem to recalibrate often, but still seemed to be similar to the previous values for the same device. Whereas if the device is totally different, the values were quite a bit different</p>",
      "rawMarkdown": "How different is each one? The android sensors seem to recalibrate often, but still seemed to be similar to the previous values for the same device. Whereas if the device is totally different, the values were quite a bit different",
      "votes": null
    },
    {
      "id": "1284421",
      "postDate": "04/25/2021 22:24:45",
      "content": "<p>The values seemed to be somewhat similar to the previous values, so it may be recalibrated. These are the stats.  </p>\n<p>max_bias_x: -28.920000076293945 <br>\nmin_bias_x: -44.279998779296875 <br>\naverage_bias_x: -35.60202447673942<br>\nmax_bias_y: 86.69999694824219 <br>\nmin_bias_y: 75.29999542236328 <br>\naverage_bias_y: 80.91607410092897<br>\nmax_bias_z: -65.63999938964844 <br>\nmin_bias_z: -86.33999633789062 <br>\naverage_bias_z: -74.5989855995661</p>",
      "rawMarkdown": "The values seemed to be somewhat similar to the previous values, so it may be recalibrated. These are the stats.  \n\nmax_bias_x: -28.920000076293945 \nmin_bias_x: -44.279998779296875 \naverage_bias_x: -35.60202447673942\nmax_bias_y: 86.69999694824219 \nmin_bias_y: 75.29999542236328 \naverage_bias_y: 80.91607410092897\nmax_bias_z: -65.63999938964844 \nmin_bias_z: -86.33999633789062 \naverage_bias_z: -74.5989855995661",
      "votes": null
    },
    {
      "id": "1284426",
      "postDate": "04/25/2021 22:41:49",
      "content": "<p>Yes, I think those would be classified as the same device with my algorithm - also check the other uncalibrated bias values, other than TYPE_MAGNETIC_FIELD (that was just for the example) - it should be fairly close for all of them, which I classify as \"drift\" on the same device</p>",
      "rawMarkdown": "Yes, I think those would be classified as the same device with my algorithm - also check the other uncalibrated bias values, other than TYPE_MAGNETIC_FIELD (that was just for the example) - it should be fairly close for all of them, which I classify as \"drift\" on the same device",
      "votes": null
    },
    {
      "id": "1285053",
      "postDate": "04/26/2021 14:26:07",
      "content": "<p>I tried to use the start/end point leakage in my Time-Series RNN model, which is somewhat similar to <a href=\"https://www.kaggle.com/ebinan92/time-series-rnn-xy-prediction\" target=\"_blank\">https://www.kaggle.com/ebinan92/time-series-rnn-xy-prediction</a>. Specifically, As the input features, I used the waypoints/floor of the start/end waypoint of other paths, whose timestamp is close to end/start waypoint of the path. However, My model didn't improve at all. I guess the reason is, as you say, this technique is effective for only some of the paths, not for all paths. Then, How did you decide which path to apply this technique to?</p>",
      "rawMarkdown": "I tried to use the start/end point leakage in my Time-Series RNN model, which is somewhat similar to https://www.kaggle.com/ebinan92/time-series-rnn-xy-prediction. Specifically, As the input features, I used the waypoints/floor of the start/end waypoint of other paths, whose timestamp is close to end/start waypoint of the path. However, My model didn't improve at all. I guess the reason is, as you say, this technique is effective for only some of the paths, not for all paths. Then, How did you decide which path to apply this technique to?",
      "votes": null
    },
    {
      "id": "1285063",
      "postDate": "04/26/2021 14:35:12",
      "content": "<p>As noted <a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/234543#1285053\" target=\"_blank\">here</a>, I tried what <a href=\"https://www.kaggle.com/tomooinubushi\" target=\"_blank\">@tomooinubushi</a> said, but my model didn't improve. </p>\n<pre><code>If you use these information as input of path-wise time series models like lstm, for example, the leaked information would propagate to the predictions of second and third waypoints and might cause broader changes.\n</code></pre>\n<p>I think this is because these information are helpful for only some paths, not for all paths. How did you decide 176 start/end waypoints? I'm not using pseudo-device/user information now, but even if I use them, I think this technique is effective for only some paths, not for all the paths.</p>",
      "rawMarkdown": "As noted [here](https://www.kaggle.com/c/indoor-location-navigation/discussion/234543#1285053), I tried what @tomooinubushi said, but my model didn't improve. \n```\nIf you use these information as input of path-wise time series models like lstm, for example, the leaked information would propagate to the predictions of second and third waypoints and might cause broader changes.\n```\nI think this is because these information are helpful for only some paths, not for all paths. How did you decide 176 start/end waypoints? I'm not using pseudo-device/user information now, but even if I use them, I think this technique is effective for only some paths, not for all the paths.",
      "votes": null
    },
    {
      "id": "1285116",
      "postDate": "04/26/2021 15:34:35",
      "content": "<p>What I wrote above is just a speculation. Please do not take it seriously.<br>\n176 start/end waypoints are selected based on temporal difference between start/end point and temporally nearest end/start point. It is implemented as start_threshold and end_threshold in <a href=\"https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage\" target=\"_blank\">my postprocessing function</a>. <br>\nThese thresholds (5,500) are based on EDA of raw timestamps. I found whether start and end points are the same is related to this temporal difference (@chris62 also mentioned this point above). In addition, I tried to use raw timestamp information only, and not to use any sensor information to demonstrate leakage in this notebook. </p>",
      "rawMarkdown": "What I wrote above is just a speculation. Please do not take it seriously.\n176 start/end waypoints are selected based on temporal difference between start/end point and temporally nearest end/start point. It is implemented as start_threshold and end_threshold in [my postprocessing function](https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage). \nThese thresholds (5,500) are based on EDA of raw timestamps. I found whether start and end points are the same is related to this temporal difference (@chris62 also mentioned this point above). In addition, I tried to use raw timestamp information only, and not to use any sensor information to demonstrate leakage in this notebook.",
      "votes": null
    },
    {
      "id": "1285120",
      "postDate": "04/26/2021 15:41:10",
      "content": "<p>There might be a couple of things going on:</p>\n<p>First, if you're using it on your current top LB model, then it's very possible that you already have all the correct start/end points, since your model is so good already! :)</p>\n<p>If it wasn't that model, then I'm not sure; but I guess I'd say that I noticed a few distinct patterns based on the timestamps between paths (you can notice this in training data and (I think) test data)… I don't know how much to say, but sometimes the start/end points are identical, and sometimes they're just close together, and sometimes the timestamp is large but then the start/end points are close anyway… </p>\n<p>My guess though, is that your top LB model has extracted all of that information anyway - and the predictions from that model are already better than what you could do by looking at start/end points.</p>",
      "rawMarkdown": "There might be a couple of things going on:\n\nFirst, if you're using it on your current top LB model, then it's very possible that you already have all the correct start/end points, since your model is so good already! :)\n\nIf it wasn't that model, then I'm not sure; but I guess I'd say that I noticed a few distinct patterns based on the timestamps between paths (you can notice this in training data and (I think) test data)... I don't know how much to say, but sometimes the start/end points are identical, and sometimes they're just close together, and sometimes the timestamp is large but then the start/end points are close anyway... \n\nMy guess though, is that your top LB model has extracted all of that information anyway - and the predictions from that model are already better than what you could do by looking at start/end points.",
      "votes": null
    },
    {
      "id": "1285123",
      "postDate": "04/26/2021 15:43:03",
      "content": "<p>Thank you, your threshold technique looks great. However, as the input of RNN, I also use the timestamp difference between start/end &amp; end/start points in addition to waypoints/floor, so it's still a bit strange my RNN didn't work well. I'll check my implementation again. </p>",
      "rawMarkdown": "Thank you, your threshold technique looks great. However, as the input of RNN, I also use the timestamp difference between start/end & end/start points in addition to waypoints/floor, so it's still a bit strange my RNN didn't work well. I'll check my implementation again.",
      "votes": null
    },
    {
      "id": "1285126",
      "postDate": "04/26/2021 15:50:43",
      "content": "<p>Thank you. As you say, when applying tomoo's leak postprocessing to our current top LB model (2.093), the score didn't change (2.093). In the past, I remember it improved the score from 3.091 to 3.055, though. </p>\n<pre><code>I noticed a few distinct patterns based on the timestamps between paths (you can notice this in training data and (I think) test data)… I don't know how much to say, but sometimes the start/end points are identical, and sometimes they're just close together, and sometimes the timestamp is large but then the start/end points are close anyway…\n</code></pre>\n<p>Hmm, seems there are some facts We didn't notice. I will try to find these patterns. </p>",
      "rawMarkdown": "Thank you. As you say, when applying tomoo's leak postprocessing to our current top LB model (2.093), the score didn't change (2.093). In the past, I remember it improved the score from 3.091 to 3.055, though. \n```\nI noticed a few distinct patterns based on the timestamps between paths (you can notice this in training data and (I think) test data)… I don't know how much to say, but sometimes the start/end points are identical, and sometimes they're just close together, and sometimes the timestamp is large but then the start/end points are close anyway…\n```\nHmm, seems there are some facts We didn't notice. I will try to find these patterns.",
      "votes": null
    },
    {
      "id": "1285132",
      "postDate": "04/26/2021 15:55:46",
      "content": "<p>Ah yes, then I would guess that your 2.093 model already has mostly correct start/end points anyway, so there's little or no benefit.</p>\n<p>I used this leak when I was at 3.5 and saw good results - so that's quite a bit higher</p>",
      "rawMarkdown": "Ah yes, then I would guess that your 2.093 model already has mostly correct start/end points anyway, so there's little or no benefit.\n\nI used this leak when I was at 3.5 and saw good results - so that's quite a bit higher",
      "votes": null
    },
    {
      "id": "1288355",
      "postDate": "04/29/2021 22:46:14",
      "content": "<p>About our 1.881 sub: <br>\nIt was achieved by solving some overfit problems.<br>\nI'm now trying to use the start/end point leakage as chris said, but still not succeeded.</p>",
      "rawMarkdown": "About our 1.881 sub: \nIt was achieved by solving some overfit problems.\nI'm now trying to use the start/end point leakage as chris said, but still not succeeded.",
      "votes": null
    },
    {
      "id": "1288370",
      "postDate": "04/29/2021 23:47:40",
      "content": "<p>1.88! 🎉</p>\n<p>I think at that point, using this data leak probably won't help much 😄; it seems like you mostly have the paths in the right spots more or less, and it's just getting all the waypoints in place now :)</p>",
      "rawMarkdown": "1.88! 🎉\n\nI think at that point, using this data leak probably won't help much 😄; it seems like you mostly have the paths in the right spots more or less, and it's just getting all the waypoints in place now :)",
      "votes": null
    },
    {
      "id": "1288375",
      "postDate": "04/30/2021 00:00:53",
      "content": "<p>I guess it won't help as a postprocessing because tomoo's kernel didn't work for us, but I guess it may be useful when used in earlier steps. I don't give up looking for the things you notice but we don't notice!</p>",
      "rawMarkdown": "I guess it won't help as a postprocessing because tomoo's kernel didn't work for us, but I guess it may be useful when used in earlier steps. I don't give up looking for the things you notice but we don't notice!",
      "votes": null
    },
    {
      "id": "1289109",
      "postDate": "04/30/2021 16:39:52",
      "content": "<p>Great work! <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> Could you please share your local validation score? Thank you.</p>",
      "rawMarkdown": "Great work! @mamasinkgs Could you please share your local validation score? Thank you.",
      "votes": null
    },
    {
      "id": "1289298",
      "postDate": "04/30/2021 20:10:25",
      "content": "<p>LB-CV gap was about 0.5 ~ 1.0 in the past, but we fixed some problems of our method and it's about 1.0 ~ 1.5 now. It is basically correlated.</p>",
      "rawMarkdown": "LB-CV gap was about 0.5 ~ 1.0 in the past, but we fixed some problems of our method and it's about 1.0 ~ 1.5 now. It is basically correlated.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1283397,
      "author_name": "higepon",
      "author_url": "",
      "post_date": "04/24/2021 22:17:18",
      "content": "<p>Thanks for sharing and being generous and fair!<br>\nI'll try to utilize the device leakage to see if it works for us.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1283493,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "04/25/2021 01:32:27",
          "content": "<p><a href=\"https://www.kaggle.com/chris62\" target=\"_blank\">@chris62</a> <br>\nJust FYI and I guess you are aware of it but there seems multiple pseudo device ids in one path file.</p>\n<pre><code>% grep TYPE_MAGNETIC_FIELD_UNCALIBRATED  5dcf888d878f3300066c6e42.txt | cut -f6-8 | uniq\n-2.3880005      1.7684937       -316.80603\n-5.015564       3.1524658       -391.75568\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1283500,
          "author_name": "chris62",
          "author_url": "",
          "post_date": "04/25/2021 01:40:34",
          "content": "<p>Yes, if the device is recalibrated mid-path, then the bias terms will change.  I don't remember seeing a drift that large, but guess I missed it :), but yes, it will \"drift\" or change over time for the same device, so I had to accept some % difference as still the same device (though totally separate devices are generally easy to separate, which is where I think it's probably most useful)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1283503,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "04/25/2021 01:44:20",
          "content": "<p>Thanks. The \"drift\" is not happening that often so I think it's should be fine :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1284411,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "04/25/2021 22:08:30",
          "content": "<p>I found 5e034414ae83d40006039063 has 158 pseudo device ids 👀</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1284414,
          "author_name": "chris62",
          "author_url": "",
          "post_date": "04/25/2021 22:14:02",
          "content": "<p>How different is each one? The android sensors seem to recalibrate often, but still seemed to be similar to the previous values for the same device. Whereas if the device is totally different, the values were quite a bit different</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1284421,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "04/25/2021 22:24:45",
          "content": "<p>The values seemed to be somewhat similar to the previous values, so it may be recalibrated. These are the stats.  </p>\n<p>max_bias_x: -28.920000076293945 <br>\nmin_bias_x: -44.279998779296875 <br>\naverage_bias_x: -35.60202447673942<br>\nmax_bias_y: 86.69999694824219 <br>\nmin_bias_y: 75.29999542236328 <br>\naverage_bias_y: 80.91607410092897<br>\nmax_bias_z: -65.63999938964844 <br>\nmin_bias_z: -86.33999633789062 <br>\naverage_bias_z: -74.5989855995661</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1284426,
          "author_name": "chris62",
          "author_url": "",
          "post_date": "04/25/2021 22:41:49",
          "content": "<p>Yes, I think those would be classified as the same device with my algorithm - also check the other uncalibrated bias values, other than TYPE_MAGNETIC_FIELD (that was just for the example) - it should be fairly close for all of them, which I classify as \"drift\" on the same device</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1283438,
      "author_name": "dehokanta",
      "author_url": "",
      "post_date": "04/24/2021 23:36:45",
      "content": "<p>I learned something new from this discussion.I would like to reflect them in my prediction.<br>\nThank you for sharing leaks with us.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1283442,
      "author_name": "tomooinubushi",
      "author_url": "",
      "post_date": "04/24/2021 23:40:22",
      "content": "<p>Thank you for sharing them. I did not notice 3rd point.<br>\nWould you elaborate why you think the last 3 float values in uncalibrated sensor information are device specific?<br>\nAre they consistent with waypoint leakage (2nd point) and/or <a href=\"https://www.kaggle.com/tomooinubushi/retrieving-user-id-from-leaked-wifi-feature\" target=\"_blank\">shared wifi records leakage</a>?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1283447,
          "author_name": "chris62",
          "author_url": "",
          "post_date": "04/24/2021 23:47:07",
          "content": "<p>The uncalibrated sensors are supplied through Android's \"uncalibrated sensors\" api: <a href=\"https://source.android.com/devices/sensors/sensor-types#uncalibrated_sensors\" target=\"_blank\">https://source.android.com/devices/sensors/sensor-types#uncalibrated_sensors</a></p>\n<p>The \"bias\" terms (the second set of 3 values) are set through the device calibration; so they are the same (or similar) for each device (though they drift over time).</p>\n<p>That doesn't mean they're identical for each device - but if you look at two different devices, they will probably be quite different; so if they are similar across two timestamps, it can be assumed to be the same device.</p>\n<p>Hope that helps!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1283452,
          "author_name": "tomooinubushi",
          "author_url": "",
          "post_date": "04/24/2021 23:56:41",
          "content": "<p>Thank you for your reply. I agree with your conclusion. <br>\nI hope this is the last leakage we can find in this competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1283470,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "04/25/2021 00:47:05",
      "content": "<p>Thank you so much for posting this. Regarding this point:</p>\n<blockquote>\n  <p>How I'm using it: this data leak is what really caused me to write this post; noticing and incorporating this is what dropped me from 3.x+ to 2.x. I didn't feel good about how well that worked, so I'm writing this post to let others know about the data leak.</p>\n</blockquote>\n<p>I remember the <a href=\"https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage\" target=\"_blank\">kernel</a>, which uses the similar idea of leak 1)+2), mentioned that the improvement was 0.1. Do you mean it can actually be 1.0? Thank you!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1283475,
          "author_name": "chris62",
          "author_url": "",
          "post_date": "04/25/2021 00:51:46",
          "content": "<p>If you just assign all the start waypoints to the previous end waypoint, and all the end waypoints to the next start waypoints, then you'll do as much harm as good, and won't get much boost :)</p>\n<p>But if you are more careful with when and how you use this info, then you can do better than 0.1</p>\n<p>It's difficult to say in my case because I made several improvements all at once - but yes, this is one of the factors I credit with my results going from 3.5 to 2.5. Perhaps it was half of that difference? (so about 0.5?) - difficult to say.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1283516,
          "author_name": "tomooinubushi",
          "author_url": "",
          "post_date": "04/25/2021 02:19:33",
          "content": "<p>I agree with <a href=\"https://www.kaggle.com/chris62\" target=\"_blank\">@chris62</a>. I should also add that my postprocessing only processed 176 start/end waypoints of 10,133 test waypoints (1.7%) and got 0.07 gain.<br>\nIf you use these information as input of path-wise time series models like lstm, for example, the leaked information would propagate to the predictions of second and third waypoints and might cause broader changes. <br>\nOne more thing I should add is I did not include device/user wise information to estimate leaked start/end waypoints in that kernel. I thought nobody use this postprocessing as is, the fact behind it is important. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1283737,
          "author_name": "jiweiliu",
          "author_url": "",
          "post_date": "04/25/2021 08:03:02",
          "content": "<p>Thank you both. Very helpful.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1285063,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "04/26/2021 14:35:12",
          "content": "<p>As noted <a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/234543#1285053\" target=\"_blank\">here</a>, I tried what <a href=\"https://www.kaggle.com/tomooinubushi\" target=\"_blank\">@tomooinubushi</a> said, but my model didn't improve. </p>\n<pre><code>If you use these information as input of path-wise time series models like lstm, for example, the leaked information would propagate to the predictions of second and third waypoints and might cause broader changes.\n</code></pre>\n<p>I think this is because these information are helpful for only some paths, not for all paths. How did you decide 176 start/end waypoints? I'm not using pseudo-device/user information now, but even if I use them, I think this technique is effective for only some paths, not for all the paths.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1285116,
          "author_name": "tomooinubushi",
          "author_url": "",
          "post_date": "04/26/2021 15:34:35",
          "content": "<p>What I wrote above is just a speculation. Please do not take it seriously.<br>\n176 start/end waypoints are selected based on temporal difference between start/end point and temporally nearest end/start point. It is implemented as start_threshold and end_threshold in <a href=\"https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage\" target=\"_blank\">my postprocessing function</a>. <br>\nThese thresholds (5,500) are based on EDA of raw timestamps. I found whether start and end points are the same is related to this temporal difference (@chris62 also mentioned this point above). In addition, I tried to use raw timestamp information only, and not to use any sensor information to demonstrate leakage in this notebook. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1285123,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "04/26/2021 15:43:03",
          "content": "<p>Thank you, your threshold technique looks great. However, as the input of RNN, I also use the timestamp difference between start/end &amp; end/start points in addition to waypoints/floor, so it's still a bit strange my RNN didn't work well. I'll check my implementation again. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1283564,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "04/25/2021 03:55:56",
      "content": "<p>Thank you for sharing. As far as I know, Our team doesn't use raw timestamp information (#1) and device information (#3) now. Youri tried to use the start/end point leakage (#2) by <a href=\"https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage\" target=\"_blank\">tomoo's kernel</a> 4 days ago, but he said the model improved from 2.715 to 2.712 (only 0.003 improvement). <br>\nSo, We don't use any leakage explicitly now. However, as <a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/231311\" target=\"_blank\">tomoo said</a>, probably we all already used this leakage by wifi. This is probably because the improvement is small.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1288355,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "04/29/2021 22:46:14",
          "content": "<p>About our 1.881 sub: <br>\nIt was achieved by solving some overfit problems.<br>\nI'm now trying to use the start/end point leakage as chris said, but still not succeeded.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1288370,
          "author_name": "chris62",
          "author_url": "",
          "post_date": "04/29/2021 23:47:40",
          "content": "<p>1.88! 🎉</p>\n<p>I think at that point, using this data leak probably won't help much 😄; it seems like you mostly have the paths in the right spots more or less, and it's just getting all the waypoints in place now :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1288375,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "04/30/2021 00:00:53",
          "content": "<p>I guess it won't help as a postprocessing because tomoo's kernel didn't work for us, but I guess it may be useful when used in earlier steps. I don't give up looking for the things you notice but we don't notice!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1289109,
          "author_name": "jiweiliu",
          "author_url": "",
          "post_date": "04/30/2021 16:39:52",
          "content": "<p>Great work! <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> Could you please share your local validation score? Thank you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1289298,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "04/30/2021 20:10:25",
          "content": "<p>LB-CV gap was about 0.5 ~ 1.0 in the past, but we fixed some problems of our method and it's about 1.0 ~ 1.5 now. It is basically correlated.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1285053,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "04/26/2021 14:26:07",
      "content": "<p>I tried to use the start/end point leakage in my Time-Series RNN model, which is somewhat similar to <a href=\"https://www.kaggle.com/ebinan92/time-series-rnn-xy-prediction\" target=\"_blank\">https://www.kaggle.com/ebinan92/time-series-rnn-xy-prediction</a>. Specifically, As the input features, I used the waypoints/floor of the start/end waypoint of other paths, whose timestamp is close to end/start waypoint of the path. However, My model didn't improve at all. I guess the reason is, as you say, this technique is effective for only some of the paths, not for all paths. Then, How did you decide which path to apply this technique to?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1285120,
          "author_name": "chris62",
          "author_url": "",
          "post_date": "04/26/2021 15:41:10",
          "content": "<p>There might be a couple of things going on:</p>\n<p>First, if you're using it on your current top LB model, then it's very possible that you already have all the correct start/end points, since your model is so good already! :)</p>\n<p>If it wasn't that model, then I'm not sure; but I guess I'd say that I noticed a few distinct patterns based on the timestamps between paths (you can notice this in training data and (I think) test data)… I don't know how much to say, but sometimes the start/end points are identical, and sometimes they're just close together, and sometimes the timestamp is large but then the start/end points are close anyway… </p>\n<p>My guess though, is that your top LB model has extracted all of that information anyway - and the predictions from that model are already better than what you could do by looking at start/end points.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1285126,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "04/26/2021 15:50:43",
          "content": "<p>Thank you. As you say, when applying tomoo's leak postprocessing to our current top LB model (2.093), the score didn't change (2.093). In the past, I remember it improved the score from 3.091 to 3.055, though. </p>\n<pre><code>I noticed a few distinct patterns based on the timestamps between paths (you can notice this in training data and (I think) test data)… I don't know how much to say, but sometimes the start/end points are identical, and sometimes they're just close together, and sometimes the timestamp is large but then the start/end points are close anyway…\n</code></pre>\n<p>Hmm, seems there are some facts We didn't notice. I will try to find these patterns. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1285132,
          "author_name": "chris62",
          "author_url": "",
          "post_date": "04/26/2021 15:55:46",
          "content": "<p>Ah yes, then I would guess that your 2.093 model already has mostly correct start/end points anyway, so there's little or no benefit.</p>\n<p>I used this leak when I was at 3.5 and saw good results - so that's quite a bit higher</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1283355": "There have been previous posts about some of these data leaks, but in order to level the playing field, I wanted to outline the 3 data leaks I know about, and how I'm using them.\n\nI don't share this information lightly, but I think it's important that the competition be fair, and I don't want unintentional data leaks to be the reason for victory :)\n\nI am sorry if some of the top teams are using these leaks secretly - but hopefully you understand why I'm publishing them.\n\n\n**1. Timestamp leakage**\n\nThe event timestamps have been reset in the test data files, but the _wifi last seen timetstamps have not_\n\nIf you look at a TYPE_WIFI record from a test path:\n\n`0000000001900 TYPE_WIFI 09f836894fc1fe9af6f429fc24dcccc2e6847fe0  b773f508054e9d212d9ef69fea73222c20eece4c  -52 2462  1573797536775`\n\nThe last value there is the true timestamp of the \"last seen\" wifi entry, even though the first value is relative to the start of the path.\n\nThis has been already reported:\n\nhttps://www.kaggle.com/c/indoor-location-navigation/discussion/228898\n\nhttps://www.kaggle.com/c/indoor-location-navigation/discussion/232072\n\nThat means that *you can slot all the test paths into the time where they were collected* relative to the training paths.\n\n*How I'm using it:* I was very excited about this at first - but honestly, it isn't as big as it seems at first, since you can get about 99% of the way there with a good model anyway. However, the biggest place this can be used is for *floor level calculations*.\n\nThe test paths are taken from a sampling of all gathered paths, and often, those test paths are taken right from the middle of a large group of training paths that are all on the same floor.\n\nNow - a simple wifi model should be able to capture 99% of the floors correctly anyway.\n\nHere's one specific example from my test set though:\n\nFor one test path, my model was predicting F4, but based on the timestamp when it occurred, it was very clear that the path _must_ have come from F3.\n\nIt is fairly simple (once you see some of these paths) to create a few simple rules about what floor it must have come from, based on which floor the surrounding paths are on.\n\n*A word of caution:* don't assume that all test paths are located on the training floors that surround them, and don't assume that only one device was collecting information at any one time!\n\n\n**2. End point is the same as some start points**\n\nThis has already been reported in a few posts:\n\nhttps://www.kaggle.com/c/indoor-location-navigation/discussion/228898\n\nhttps://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage\n\nThe idea is: some paths have the exact same stop/end waypoints as the previous or next paths.\n\nIf you combine that with my leak #1, that means that for some posts, you can know exactly what the start or stop points are.\n\nA few notes about this data leak: A) it only applies to some of the test paths, but B) if you generally use that idea combined with leak #1, that test paths close in time to training paths will generally be in similar areas of a site/floor - then you can sometimes correlate the end or start of training paths with the start or end of test paths...\n\nSome of that should be obvious: if one person is recording both training and test paths, then if they only have a few seconds between paths, they could have only traveled a few meters...\n\n**How I'm using it**: this data leak is what really caused me to write this post; noticing and incorporating this is what dropped me from 3.x+ to 2.x. I didn't feel good about how well that worked, so I'm writing this post to let others know about the data leak.\n\n\n**3. Device leakage**\n\nThis is the data leak that I don't think has been talked about on this forum before.\n\nIf you look at the uncalibrated sensor information, you might see a row like:\n\n`0000000000276 TYPE_MAGNETIC_FIELD_UNCALIBRATED  -106.762695 35.142517 -355.44434  -83.41217 16.18042  -325.82092  3`\n\nThose last 3 float values (before the accuracy) are **device specific**. Since the same devices are used for the training and test paths, it SEEMS like this is a huge leak - because combining with leak #1, I can almost exactly specify what device was recording which path at what time. (and so exactly specify the test floors)\n\nAgain, this seems like a huge leakage, but when combined with #1 and #2, is actually secondary... but can still be very helpful as a meta feature to certain models.\n\n**How I'm using this**: I'm using this \"device\" feature in my models, so I'm a bit unsure about how little or much this is influencing my results. If you are using the uncalibrated sensor values, then you should see about as much improvement as I'm seeing anyway though.\n\nI know this can help a bit with floor level predictions - but again, if you have a generally good wifi/beacon model, then this data leak probably doesn't matter that much.\n\n\n# Why post this\n\nHopefully this discussion post has provided some insight about the competition... again - I'm sorry if some of the top teams were using these leaks in secret, but I think that it is a much more fair competition if the leaks are made public.\n\nIn my own models, these leaks have undoubtedly provided some boost - but I'd really like to see the winning models without the leak advantages - so hopefully this helps all teams.\n\n\n\nThanks,\nChris",
    "1283397": "Thanks for sharing and being generous and fair!\nI'll try to utilize the device leakage to see if it works for us.",
    "1283438": "I learned something new from this discussion.I would like to reflect them in my prediction.\nThank you for sharing leaks with us.",
    "1283442": "Thank you for sharing them. I did not notice 3rd point.\nWould you elaborate why you think the last 3 float values in uncalibrated sensor information are device specific?\nAre they consistent with waypoint leakage (2nd point) and/or [shared wifi records leakage](https://www.kaggle.com/tomooinubushi/retrieving-user-id-from-leaked-wifi-feature)?",
    "1283447": "The uncalibrated sensors are supplied through Android's \"uncalibrated sensors\" api: https://source.android.com/devices/sensors/sensor-types#uncalibrated_sensors\n\nThe \"bias\" terms (the second set of 3 values) are set through the device calibration; so they are the same (or similar) for each device (though they drift over time).\n\nThat doesn't mean they're identical for each device - but if you look at two different devices, they will probably be quite different; so if they are similar across two timestamps, it can be assumed to be the same device.\n\nHope that helps!",
    "1283452": "Thank you for your reply. I agree with your conclusion. \nI hope this is the last leakage we can find in this competition.",
    "1283470": "Thank you so much for posting this. Regarding this point:\n>How I'm using it: this data leak is what really caused me to write this post; noticing and incorporating this is what dropped me from 3.x+ to 2.x. I didn't feel good about how well that worked, so I'm writing this post to let others know about the data leak.\n\nI remember the [kernel](https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage), which uses the similar idea of leak 1)+2), mentioned that the improvement was 0.1. Do you mean it can actually be 1.0? Thank you!",
    "1283475": "If you just assign all the start waypoints to the previous end waypoint, and all the end waypoints to the next start waypoints, then you'll do as much harm as good, and won't get much boost :)\n\nBut if you are more careful with when and how you use this info, then you can do better than 0.1\n\nIt's difficult to say in my case because I made several improvements all at once - but yes, this is one of the factors I credit with my results going from 3.5 to 2.5. Perhaps it was half of that difference? (so about 0.5?) - difficult to say.",
    "1283493": "chris62 \nJust FYI and I guess you are aware of it but there seems multiple pseudo device ids in one path file.\n\n```\n% grep TYPE_MAGNETIC_FIELD_UNCALIBRATED  5dcf888d878f3300066c6e42.txt | cut -f6-8 | uniq\n-2.3880005      1.7684937       -316.80603\n-5.015564       3.1524658       -391.75568\n```",
    "1283500": "Yes, if the device is recalibrated mid-path, then the bias terms will change.  I don't remember seeing a drift that large, but guess I missed it :), but yes, it will \"drift\" or change over time for the same device, so I had to accept some % difference as still the same device (though totally separate devices are generally easy to separate, which is where I think it's probably most useful)",
    "1283503": "Thanks. The \"drift\" is not happening that often so I think it's should be fine :)",
    "1283516": "I agree with @chris62. I should also add that my postprocessing only processed 176 start/end waypoints of 10,133 test waypoints (1.7%) and got 0.07 gain.\nIf you use these information as input of path-wise time series models like lstm, for example, the leaked information would propagate to the predictions of second and third waypoints and might cause broader changes. \nOne more thing I should add is I did not include device/user wise information to estimate leaked start/end waypoints in that kernel. I thought nobody use this postprocessing as is, the fact behind it is important.",
    "1283564": "Thank you for sharing. As far as I know, Our team doesn't use raw timestamp information (#1) and device information (#3) now. Youri tried to use the start/end point leakage (#2) by [tomoo's kernel](https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage) 4 days ago, but he said the model improved from 2.715 to 2.712 (only 0.003 improvement). \nSo, We don't use any leakage explicitly now. However, as [tomoo said](https://www.kaggle.com/c/indoor-location-navigation/discussion/231311), probably we all already used this leakage by wifi. This is probably because the improvement is small.",
    "1283737": "Thank you both. Very helpful.",
    "1284411": "I found 5e034414ae83d40006039063 has 158 pseudo device ids 👀",
    "1284414": "How different is each one? The android sensors seem to recalibrate often, but still seemed to be similar to the previous values for the same device. Whereas if the device is totally different, the values were quite a bit different",
    "1284421": "The values seemed to be somewhat similar to the previous values, so it may be recalibrated. These are the stats.  \n\nmax_bias_x: -28.920000076293945 \nmin_bias_x: -44.279998779296875 \naverage_bias_x: -35.60202447673942\nmax_bias_y: 86.69999694824219 \nmin_bias_y: 75.29999542236328 \naverage_bias_y: 80.91607410092897\nmax_bias_z: -65.63999938964844 \nmin_bias_z: -86.33999633789062 \naverage_bias_z: -74.5989855995661",
    "1284426": "Yes, I think those would be classified as the same device with my algorithm - also check the other uncalibrated bias values, other than TYPE_MAGNETIC_FIELD (that was just for the example) - it should be fairly close for all of them, which I classify as \"drift\" on the same device",
    "1285053": "I tried to use the start/end point leakage in my Time-Series RNN model, which is somewhat similar to https://www.kaggle.com/ebinan92/time-series-rnn-xy-prediction. Specifically, As the input features, I used the waypoints/floor of the start/end waypoint of other paths, whose timestamp is close to end/start waypoint of the path. However, My model didn't improve at all. I guess the reason is, as you say, this technique is effective for only some of the paths, not for all paths. Then, How did you decide which path to apply this technique to?",
    "1285063": "As noted [here](https://www.kaggle.com/c/indoor-location-navigation/discussion/234543#1285053), I tried what @tomooinubushi said, but my model didn't improve. \n```\nIf you use these information as input of path-wise time series models like lstm, for example, the leaked information would propagate to the predictions of second and third waypoints and might cause broader changes.\n```\nI think this is because these information are helpful for only some paths, not for all paths. How did you decide 176 start/end waypoints? I'm not using pseudo-device/user information now, but even if I use them, I think this technique is effective for only some paths, not for all the paths.",
    "1285116": "What I wrote above is just a speculation. Please do not take it seriously.\n176 start/end waypoints are selected based on temporal difference between start/end point and temporally nearest end/start point. It is implemented as start_threshold and end_threshold in [my postprocessing function](https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage). \nThese thresholds (5,500) are based on EDA of raw timestamps. I found whether start and end points are the same is related to this temporal difference (@chris62 also mentioned this point above). In addition, I tried to use raw timestamp information only, and not to use any sensor information to demonstrate leakage in this notebook.",
    "1285120": "There might be a couple of things going on:\n\nFirst, if you're using it on your current top LB model, then it's very possible that you already have all the correct start/end points, since your model is so good already! :)\n\nIf it wasn't that model, then I'm not sure; but I guess I'd say that I noticed a few distinct patterns based on the timestamps between paths (you can notice this in training data and (I think) test data)... I don't know how much to say, but sometimes the start/end points are identical, and sometimes they're just close together, and sometimes the timestamp is large but then the start/end points are close anyway... \n\nMy guess though, is that your top LB model has extracted all of that information anyway - and the predictions from that model are already better than what you could do by looking at start/end points.",
    "1285123": "Thank you, your threshold technique looks great. However, as the input of RNN, I also use the timestamp difference between start/end & end/start points in addition to waypoints/floor, so it's still a bit strange my RNN didn't work well. I'll check my implementation again.",
    "1285126": "Thank you. As you say, when applying tomoo's leak postprocessing to our current top LB model (2.093), the score didn't change (2.093). In the past, I remember it improved the score from 3.091 to 3.055, though. \n```\nI noticed a few distinct patterns based on the timestamps between paths (you can notice this in training data and (I think) test data)… I don't know how much to say, but sometimes the start/end points are identical, and sometimes they're just close together, and sometimes the timestamp is large but then the start/end points are close anyway…\n```\nHmm, seems there are some facts We didn't notice. I will try to find these patterns.",
    "1285132": "Ah yes, then I would guess that your 2.093 model already has mostly correct start/end points anyway, so there's little or no benefit.\n\nI used this leak when I was at 3.5 and saw good results - so that's quite a bit higher",
    "1288355": "About our 1.881 sub: \nIt was achieved by solving some overfit problems.\nI'm now trying to use the start/end point leakage as chris said, but still not succeeded.",
    "1288370": "1.88! 🎉\n\nI think at that point, using this data leak probably won't help much 😄; it seems like you mostly have the paths in the right spots more or less, and it's just getting all the waypoints in place now :)",
    "1288375": "I guess it won't help as a postprocessing because tomoo's kernel didn't work for us, but I guess it may be useful when used in earlier steps. I don't give up looking for the things you notice but we don't notice!",
    "1289109": "Great work! @mamasinkgs Could you please share your local validation score? Thank you.",
    "1289298": "LB-CV gap was about 0.5 ~ 1.0 in the past, but we fixed some problems of our method and it's about 1.0 ~ 1.5 now. It is basically correlated."
  },
  "source": "meta"
}