{
  "id": 236073,
  "title": "Simple trick to check whether your floor submission is perfect in public LB",
  "url": "/competitions/indoor-location-navigation/discussion/236073",
  "author_name": "",
  "post_date": "2021-05-02T17:41:22.537144200Z",
  "votes": 24,
  "comment_count": 18,
  "views": 0,
  "content": "<p>This trick uses only 2 submissions. </p>\n<pre><code>sub = pd.read_csv('original_submission.csv')\nsub['floor'] = sub['floor'] + 1\n# submit this\nsub.to_csv('floor_plus_1.csv', index = False)\n\nsub = pd.read_csv('original_submission.csv')\nsub['floor'] = sub['floor'] - 1\n# submit this\nsub.to_csv('floor_minus_1.csv', index = False)\n</code></pre>\n<p>Then, if the LB scores of <code>floor_plus_1.csv</code> and <code>floor_minus_1.csv</code> is higher than <code>original_submission.csv</code> by 15, it means your submission is perfect in public LB, unless your submission contains big mistakes (e.g. the prediction is F3 but the target is F1) . </p>\n<p>I noticed this trick by <a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/235328#1288087\" target=\"_blank\">this comment</a>, thanks <a href=\"https://www.kaggle.com/jiweiliu\" target=\"_blank\">@jiweiliu</a> 👍</p>\n<p>By using this trick, I think we can check whether the floor predictions of <a href=\"https://www.kaggle.com/nigelhenry/simple-99-accurate-floor-model\" target=\"_blank\">simple 99% accurate floor model</a> and <a href=\"https://www.kaggle.com/jwilliamhughdore/99-80-floor-accurate-model-blstm\" target=\"_blank\">99.80% accurate floor model</a> are perfect. I hope someone to check this and report the scores of them 🙏</p>",
  "messages": [
    {
      "id": "1291108",
      "postDate": "05/02/2021 17:41:22",
      "content": "<p>This trick uses only 2 submissions. </p>\n<pre><code>sub = pd.read_csv('original_submission.csv')\nsub['floor'] = sub['floor'] + 1\n# submit this\nsub.to_csv('floor_plus_1.csv', index = False)\n\nsub = pd.read_csv('original_submission.csv')\nsub['floor'] = sub['floor'] - 1\n# submit this\nsub.to_csv('floor_minus_1.csv', index = False)\n</code></pre>\n<p>Then, if the LB scores of <code>floor_plus_1.csv</code> and <code>floor_minus_1.csv</code> is higher than <code>original_submission.csv</code> by 15, it means your submission is perfect in public LB, unless your submission contains big mistakes (e.g. the prediction is F3 but the target is F1) . </p>\n<p>I noticed this trick by <a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/235328#1288087\" target=\"_blank\">this comment</a>, thanks <a href=\"https://www.kaggle.com/jiweiliu\" target=\"_blank\">@jiweiliu</a> 👍</p>\n<p>By using this trick, I think we can check whether the floor predictions of <a href=\"https://www.kaggle.com/nigelhenry/simple-99-accurate-floor-model\" target=\"_blank\">simple 99% accurate floor model</a> and <a href=\"https://www.kaggle.com/jwilliamhughdore/99-80-floor-accurate-model-blstm\" target=\"_blank\">99.80% accurate floor model</a> are perfect. I hope someone to check this and report the scores of them 🙏</p>",
      "rawMarkdown": "This trick uses only 2 submissions. \n```\nsub = pd.read_csv('original_submission.csv')\nsub['floor'] = sub['floor'] + 1\n# submit this\nsub.to_csv('floor_plus_1.csv', index = False)\n\nsub = pd.read_csv('original_submission.csv')\nsub['floor'] = sub['floor'] - 1\n# submit this\nsub.to_csv('floor_minus_1.csv', index = False)\n```\nThen, if the LB scores of `floor_plus_1.csv` and `floor_minus_1.csv` is higher than `original_submission.csv` by 15, it means your submission is perfect in public LB, unless your submission contains big mistakes (e.g. the prediction is F3 but the target is F1) . \n\nI noticed this trick by [this comment](https://www.kaggle.com/c/indoor-location-navigation/discussion/235328#1288087), thanks @jiweiliu 👍\n\nBy using this trick, I think we can check whether the floor predictions of [simple 99% accurate floor model](https://www.kaggle.com/nigelhenry/simple-99-accurate-floor-model) and [99.80% accurate floor model](https://www.kaggle.com/jwilliamhughdore/99-80-floor-accurate-model-blstm) are perfect. I hope someone to check this and report the scores of them 🙏",
      "votes": null
    },
    {
      "id": "1291114",
      "postDate": "05/02/2021 17:45:35",
      "content": "<p>In addition, as <a href=\"https://www.kaggle.com/jiweiliu\" target=\"_blank\">@jiweiliu</a> said, we can use this trick to hide public LB score :)</p>",
      "rawMarkdown": "In addition, as @jiweiliu said, we can use this trick to hide public LB score :)",
      "votes": null
    },
    {
      "id": "1291122",
      "postDate": "05/02/2021 17:49:35",
      "content": "<p>Thanks for the trick <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> </p>",
      "rawMarkdown": "Thanks for the trick @mamasinkgs",
      "votes": null
    },
    {
      "id": "1292307",
      "postDate": "05/03/2021 20:17:18",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a>, I indeed seem to have them all correct. Since all my public LB floor predictions are identical to <a href=\"https://www.kaggle.com/nigelhenry\" target=\"_blank\">@nigelhenry</a>'s \"99% accurate\" model, we can be fairly sure that those too are 100% accurate for the public LB.</p>",
      "rawMarkdown": "Thanks @mamasinkgs, I indeed seem to have them all correct. Since all my public LB floor predictions are identical to @nigelhenry's \"99% accurate\" model, we can be fairly sure that those too are 100% accurate for the public LB.",
      "votes": null
    },
    {
      "id": "1292311",
      "postDate": "05/03/2021 20:21:51",
      "content": "<p>Thanks, but how did you check your public LB floor predictions are same as them? Do you know what samples are in public LB?</p>",
      "rawMarkdown": "Thanks, but how did you check your public LB floor predictions are same as them? Do you know what samples are in public LB?",
      "votes": null
    },
    {
      "id": "1292317",
      "postDate": "05/03/2021 20:29:23",
      "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a>, I have interchanged these two sets of floor predictions before and they always obtain the same score (obviously with identical xy coordinates).</p>",
      "rawMarkdown": "mamasinkgs, I have interchanged these two sets of floor predictions before and they always obtain the same score (obviously with identical xy coordinates).",
      "votes": null
    },
    {
      "id": "1292390",
      "postDate": "05/03/2021 22:27:07",
      "content": "<p>Thanks, got it. So, as you say, the scores of the public kernels are perfect in public LB.</p>",
      "rawMarkdown": "Thanks, got it. So, as you say, the scores of the public kernels are perfect in public LB.",
      "votes": null
    },
    {
      "id": "1292859",
      "postDate": "05/04/2021 10:54:16",
      "content": "<p>The 99% one - yes I can vouch for that (but only in public LB). I'll leave it for someone else to test the 99.8% one.</p>",
      "rawMarkdown": "The 99% one - yes I can vouch for that (but only in public LB). I'll leave it for someone else to test the 99.8% one.",
      "votes": null
    },
    {
      "id": "1294402",
      "postDate": "05/05/2021 16:08:19",
      "content": "<p>The downside of publishing a public notebook with no errors 2 months ago is that there have been very few public floor prediction models since then.</p>",
      "rawMarkdown": "The downside of publishing a public notebook with no errors 2 months ago is that there have been very few public floor prediction models since then.",
      "votes": null
    },
    {
      "id": "1294981",
      "postDate": "05/06/2021 05:25:54",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a>,<br>\nbased on an earlier discussion I was adapting <a href=\"https://www.kaggle.com/nigelhenry\" target=\"_blank\">@nigelhenry</a>'s  \"99% accurate\" model and there were around 225 floors for which I had a difference. I did not change the x y coords, only the floor<br>\nthe public LB score was identical -- no change at all !!</p>\n<p>Either (a) the 225 floors are in the hidden set (85% not counted for Public LB) or (b) the errors set themselves off. some were higher by one and some lower (vs ground truth) and the net result was same in Public LB</p>\n<p>how do you decide which floor model to use, given ground truth is not known..</p>",
      "rawMarkdown": "Hi @mamasinkgs,\nbased on an earlier discussion I was adapting @nigelhenry's  \"99% accurate\" model and there were around 225 floors for which I had a difference. I did not change the x y coords, only the floor\nthe public LB score was identical -- no change at all !!\n\nEither (a) the 225 floors are in the hidden set (85% not counted for Public LB) or (b) the errors set themselves off. some were higher by one and some lower (vs ground truth) and the net result was same in Public LB\n\nhow do you decide which floor model to use, given ground truth is not known..",
      "votes": null
    },
    {
      "id": "1295003",
      "postDate": "05/06/2021 06:10:27",
      "content": "<p>I guess the public/private split is groupkfold (group = path), so it is not strange the 225 floors are all in the private LB. So, I guess the answer is (a). To decide which floor model to use is difficult problem.. If you have multiple predictions, I think majority voting may work.</p>",
      "rawMarkdown": "I guess the public/private split is groupkfold (group = path), so it is not strange the 225 floors are all in the private LB. So, I guess the answer is (a). To decide which floor model to use is difficult problem.. If you have multiple predictions, I think majority voting may work.",
      "votes": null
    },
    {
      "id": "1295032",
      "postDate": "05/06/2021 06:41:54",
      "content": "<p>Hi, </p>\n<p>Just wanted to share my idea.  I created a set of unique bssids for each floor; (i.e) for each floor collect bssids which occurs in only that floor. My thinking was that certain bssids are detected in may floors and its difficult to identify the floor based on such commonly occurring bssids. Once such a dictionary is created for each floor in building, for each test path we can find number of bssids matching for each floor data in dictionary. Once such table is created, we can compute data such as count, mean, median of each floor's distribution and select the floor. </p>\n<p>I have implemented this idea in this <a href=\"https://www.kaggle.com/suryajrrafl/wifi-based-floor-mapping/data?scriptVersionId=60842281\" target=\"_blank\">notebook</a>. 3rd version is my solution, 4th version is the public solution by nigel. The approach I took has 100%accuracy in training set but score's poorly compared to <a href=\"https://www.kaggle.com/nigelhenry/simple-99-accurate-floor-model\" target=\"_blank\">nigel's solution</a> and found many differences. </p>\n<p>Any feedback on the approach is most welcome.</p>",
      "rawMarkdown": "Hi, \n\nJust wanted to share my idea.  I created a set of unique bssids for each floor; (i.e) for each floor collect bssids which occurs in only that floor. My thinking was that certain bssids are detected in may floors and its difficult to identify the floor based on such commonly occurring bssids. Once such a dictionary is created for each floor in building, for each test path we can find number of bssids matching for each floor data in dictionary. Once such table is created, we can compute data such as count, mean, median of each floor's distribution and select the floor. \n\n\nI have implemented this idea in this [notebook](https://www.kaggle.com/suryajrrafl/wifi-based-floor-mapping/data?scriptVersionId=60842281). 3rd version is my solution, 4th version is the public solution by nigel. The approach I took has 100%accuracy in training set but score's poorly compared to [nigel's solution](https://www.kaggle.com/nigelhenry/simple-99-accurate-floor-model) and found many differences. \n\nAny feedback on the approach is most welcome.",
      "votes": null
    },
    {
      "id": "1295201",
      "postDate": "05/06/2021 09:23:18",
      "content": "<p>aah yes. hadn't thought of majority voting!  let me try that… thanks for the suggestion <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> </p>",
      "rawMarkdown": "aah yes. hadn't thought of majority voting!  let me try that... thanks for the suggestion @mamasinkgs",
      "votes": null
    },
    {
      "id": "1295885",
      "postDate": "05/06/2021 19:12:53",
      "content": "<p><a href=\"https://www.kaggle.com/kmldas\" target=\"_blank\">@kmldas</a> similarly to you, I agree with the \"99%\" model throughout the public LB (and the 'trick' described in this current thread confirms their correctness) but disagree on 213 points comprising 11 paths in the private LB. Differences are in buildings:</p>\n<p>5d2709bb03f801723c32852c (2 paths)<br>\n5d2709d403f801723c32bd39 (3 paths)<br>\n5da138274db8ce0c98bbd3d2 (2 paths)<br>\n5da1382d4db8ce0c98bbe92e (1 path)<br>\n5da138754db8ce0c98bca82f (1 path)<br>\n5da1389e4db8ce0c98bd0547 (1 path)<br>\n5da138b74db8ce0c98bd4774 (1 path)</p>",
      "rawMarkdown": "kmldas similarly to you, I agree with the \"99%\" model throughout the public LB (and the 'trick' described in this current thread confirms their correctness) but disagree on 213 points comprising 11 paths in the private LB. Differences are in buildings:\n\n5d2709bb03f801723c32852c (2 paths)\n5d2709d403f801723c32bd39 (3 paths)\n5da138274db8ce0c98bbd3d2 (2 paths)\n5da1382d4db8ce0c98bbe92e (1 path)\n5da138754db8ce0c98bca82f (1 path)\n5da1389e4db8ce0c98bd0547 (1 path)\n5da138b74db8ce0c98bd4774 (1 path)",
      "votes": null
    },
    {
      "id": "1296186",
      "postDate": "05/07/2021 04:06:56",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/jbomitchell\" target=\"_blank\">@jbomitchell</a> .. good to know we have some differences… this may be critical in the final shakeup….hoping we land up in the right side of it 👍😄</p>",
      "rawMarkdown": "Thanks for sharing @jbomitchell .. good to know we have some differences... this may be critical in the final shakeup....hoping we land up in the right side of it 👍😄",
      "votes": null
    },
    {
      "id": "1298536",
      "postDate": "05/09/2021 01:44:52",
      "content": "<p>In addition to the floor hack, you can also estimate positional bias by adding delta_x, delta_y to your submission. With site  dependent delta (including floor), LB may show which sites are in public or not. Btw, is enough time left to try this?</p>",
      "rawMarkdown": "In addition to the floor hack, you can also estimate positional bias by adding delta_x, delta_y to your submission. With site  dependent delta (including floor), LB may show which sites are in public or not. Btw, is enough time left to try this?",
      "votes": null
    },
    {
      "id": "1298579",
      "postDate": "05/09/2021 03:14:45",
      "content": "<p>according to this <a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/224560\" target=\"_blank\">https://www.kaggle.com/c/indoor-location-navigation/discussion/224560</a> fractional floors are allowed. therefore you can hedge by submitting your (unrounded) mean of all your floor prediction models</p>",
      "rawMarkdown": "according to this https://www.kaggle.com/c/indoor-location-navigation/discussion/224560 fractional floors are allowed. therefore you can hedge by submitting your (unrounded) mean of all your floor prediction models",
      "votes": null
    },
    {
      "id": "1298594",
      "postDate": "05/09/2021 03:43:40",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/nigelhenry\" target=\"_blank\">@nigelhenry</a> … i had missed that discussion <br>\nI submitted average of the submissions over 200 were different and yes I get the same score as I think these are all in the hidden set… will know when the results are out which was the better choice… to take average or not!</p>",
      "rawMarkdown": "Thanks @nigelhenry ... i had missed that discussion \nI submitted average of the submissions over 200 were different and yes I get the same score as I think these are all in the hidden set... will know when the results are out which was the better choice... to take average or not!",
      "votes": null
    },
    {
      "id": "1301558",
      "postDate": "05/11/2021 06:21:55",
      "content": "<p>Since the floor penalty uses the absolute difference between the true floor and predicted floor, hedging by predicting the average of the floors is worse in expectation when you are &gt;50% confident of one of the floors.</p>",
      "rawMarkdown": "Since the floor penalty uses the absolute difference between the true floor and predicted floor, hedging by predicting the average of the floors is worse in expectation when you are >50% confident of one of the floors.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1291114,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "05/02/2021 17:45:35",
      "content": "<p>In addition, as <a href=\"https://www.kaggle.com/jiweiliu\" target=\"_blank\">@jiweiliu</a> said, we can use this trick to hide public LB score :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1291122,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "05/02/2021 17:49:35",
      "content": "<p>Thanks for the trick <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1292307,
      "author_name": "jbomitchell",
      "author_url": "",
      "post_date": "05/03/2021 20:17:18",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a>, I indeed seem to have them all correct. Since all my public LB floor predictions are identical to <a href=\"https://www.kaggle.com/nigelhenry\" target=\"_blank\">@nigelhenry</a>'s \"99% accurate\" model, we can be fairly sure that those too are 100% accurate for the public LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1292311,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "05/03/2021 20:21:51",
          "content": "<p>Thanks, but how did you check your public LB floor predictions are same as them? Do you know what samples are in public LB?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1292317,
          "author_name": "jbomitchell",
          "author_url": "",
          "post_date": "05/03/2021 20:29:23",
          "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a>, I have interchanged these two sets of floor predictions before and they always obtain the same score (obviously with identical xy coordinates).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1292390,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "05/03/2021 22:27:07",
          "content": "<p>Thanks, got it. So, as you say, the scores of the public kernels are perfect in public LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1292859,
          "author_name": "jbomitchell",
          "author_url": "",
          "post_date": "05/04/2021 10:54:16",
          "content": "<p>The 99% one - yes I can vouch for that (but only in public LB). I'll leave it for someone else to test the 99.8% one.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1294402,
          "author_name": "nigelhenry",
          "author_url": "",
          "post_date": "05/05/2021 16:08:19",
          "content": "<p>The downside of publishing a public notebook with no errors 2 months ago is that there have been very few public floor prediction models since then.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1294981,
      "author_name": "kmldas",
      "author_url": "",
      "post_date": "05/06/2021 05:25:54",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a>,<br>\nbased on an earlier discussion I was adapting <a href=\"https://www.kaggle.com/nigelhenry\" target=\"_blank\">@nigelhenry</a>'s  \"99% accurate\" model and there were around 225 floors for which I had a difference. I did not change the x y coords, only the floor<br>\nthe public LB score was identical -- no change at all !!</p>\n<p>Either (a) the 225 floors are in the hidden set (85% not counted for Public LB) or (b) the errors set themselves off. some were higher by one and some lower (vs ground truth) and the net result was same in Public LB</p>\n<p>how do you decide which floor model to use, given ground truth is not known..</p>",
      "votes": null,
      "replies": [
        {
          "id": 1295003,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "05/06/2021 06:10:27",
          "content": "<p>I guess the public/private split is groupkfold (group = path), so it is not strange the 225 floors are all in the private LB. So, I guess the answer is (a). To decide which floor model to use is difficult problem.. If you have multiple predictions, I think majority voting may work.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1295201,
          "author_name": "kmldas",
          "author_url": "",
          "post_date": "05/06/2021 09:23:18",
          "content": "<p>aah yes. hadn't thought of majority voting!  let me try that… thanks for the suggestion <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1295885,
          "author_name": "jbomitchell",
          "author_url": "",
          "post_date": "05/06/2021 19:12:53",
          "content": "<p><a href=\"https://www.kaggle.com/kmldas\" target=\"_blank\">@kmldas</a> similarly to you, I agree with the \"99%\" model throughout the public LB (and the 'trick' described in this current thread confirms their correctness) but disagree on 213 points comprising 11 paths in the private LB. Differences are in buildings:</p>\n<p>5d2709bb03f801723c32852c (2 paths)<br>\n5d2709d403f801723c32bd39 (3 paths)<br>\n5da138274db8ce0c98bbd3d2 (2 paths)<br>\n5da1382d4db8ce0c98bbe92e (1 path)<br>\n5da138754db8ce0c98bca82f (1 path)<br>\n5da1389e4db8ce0c98bd0547 (1 path)<br>\n5da138b74db8ce0c98bd4774 (1 path)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1296186,
          "author_name": "kmldas",
          "author_url": "",
          "post_date": "05/07/2021 04:06:56",
          "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/jbomitchell\" target=\"_blank\">@jbomitchell</a> .. good to know we have some differences… this may be critical in the final shakeup….hoping we land up in the right side of it 👍😄</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1298579,
          "author_name": "nigelhenry",
          "author_url": "",
          "post_date": "05/09/2021 03:14:45",
          "content": "<p>according to this <a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/224560\" target=\"_blank\">https://www.kaggle.com/c/indoor-location-navigation/discussion/224560</a> fractional floors are allowed. therefore you can hedge by submitting your (unrounded) mean of all your floor prediction models</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1298594,
          "author_name": "kmldas",
          "author_url": "",
          "post_date": "05/09/2021 03:43:40",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/nigelhenry\" target=\"_blank\">@nigelhenry</a> … i had missed that discussion <br>\nI submitted average of the submissions over 200 were different and yes I get the same score as I think these are all in the hidden set… will know when the results are out which was the better choice… to take average or not!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1301558,
          "author_name": "tvdwiele",
          "author_url": "",
          "post_date": "05/11/2021 06:21:55",
          "content": "<p>Since the floor penalty uses the absolute difference between the true floor and predicted floor, hedging by predicting the average of the floors is worse in expectation when you are &gt;50% confident of one of the floors.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1295032,
      "author_name": "suryajrrafl",
      "author_url": "",
      "post_date": "05/06/2021 06:41:54",
      "content": "<p>Hi, </p>\n<p>Just wanted to share my idea.  I created a set of unique bssids for each floor; (i.e) for each floor collect bssids which occurs in only that floor. My thinking was that certain bssids are detected in may floors and its difficult to identify the floor based on such commonly occurring bssids. Once such a dictionary is created for each floor in building, for each test path we can find number of bssids matching for each floor data in dictionary. Once such table is created, we can compute data such as count, mean, median of each floor's distribution and select the floor. </p>\n<p>I have implemented this idea in this <a href=\"https://www.kaggle.com/suryajrrafl/wifi-based-floor-mapping/data?scriptVersionId=60842281\" target=\"_blank\">notebook</a>. 3rd version is my solution, 4th version is the public solution by nigel. The approach I took has 100%accuracy in training set but score's poorly compared to <a href=\"https://www.kaggle.com/nigelhenry/simple-99-accurate-floor-model\" target=\"_blank\">nigel's solution</a> and found many differences. </p>\n<p>Any feedback on the approach is most welcome.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1298536,
      "author_name": "zeemeen",
      "author_url": "",
      "post_date": "05/09/2021 01:44:52",
      "content": "<p>In addition to the floor hack, you can also estimate positional bias by adding delta_x, delta_y to your submission. With site  dependent delta (including floor), LB may show which sites are in public or not. Btw, is enough time left to try this?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1291108": "This trick uses only 2 submissions. \n```\nsub = pd.read_csv('original_submission.csv')\nsub['floor'] = sub['floor'] + 1\n# submit this\nsub.to_csv('floor_plus_1.csv', index = False)\n\nsub = pd.read_csv('original_submission.csv')\nsub['floor'] = sub['floor'] - 1\n# submit this\nsub.to_csv('floor_minus_1.csv', index = False)\n```\nThen, if the LB scores of `floor_plus_1.csv` and `floor_minus_1.csv` is higher than `original_submission.csv` by 15, it means your submission is perfect in public LB, unless your submission contains big mistakes (e.g. the prediction is F3 but the target is F1) . \n\nI noticed this trick by [this comment](https://www.kaggle.com/c/indoor-location-navigation/discussion/235328#1288087), thanks @jiweiliu 👍\n\nBy using this trick, I think we can check whether the floor predictions of [simple 99% accurate floor model](https://www.kaggle.com/nigelhenry/simple-99-accurate-floor-model) and [99.80% accurate floor model](https://www.kaggle.com/jwilliamhughdore/99-80-floor-accurate-model-blstm) are perfect. I hope someone to check this and report the scores of them 🙏",
    "1291114": "In addition, as @jiweiliu said, we can use this trick to hide public LB score :)",
    "1291122": "Thanks for the trick @mamasinkgs",
    "1292307": "Thanks @mamasinkgs, I indeed seem to have them all correct. Since all my public LB floor predictions are identical to @nigelhenry's \"99% accurate\" model, we can be fairly sure that those too are 100% accurate for the public LB.",
    "1292311": "Thanks, but how did you check your public LB floor predictions are same as them? Do you know what samples are in public LB?",
    "1292317": "mamasinkgs, I have interchanged these two sets of floor predictions before and they always obtain the same score (obviously with identical xy coordinates).",
    "1292390": "Thanks, got it. So, as you say, the scores of the public kernels are perfect in public LB.",
    "1292859": "The 99% one - yes I can vouch for that (but only in public LB). I'll leave it for someone else to test the 99.8% one.",
    "1294402": "The downside of publishing a public notebook with no errors 2 months ago is that there have been very few public floor prediction models since then.",
    "1294981": "Hi @mamasinkgs,\nbased on an earlier discussion I was adapting @nigelhenry's  \"99% accurate\" model and there were around 225 floors for which I had a difference. I did not change the x y coords, only the floor\nthe public LB score was identical -- no change at all !!\n\nEither (a) the 225 floors are in the hidden set (85% not counted for Public LB) or (b) the errors set themselves off. some were higher by one and some lower (vs ground truth) and the net result was same in Public LB\n\nhow do you decide which floor model to use, given ground truth is not known..",
    "1295003": "I guess the public/private split is groupkfold (group = path), so it is not strange the 225 floors are all in the private LB. So, I guess the answer is (a). To decide which floor model to use is difficult problem.. If you have multiple predictions, I think majority voting may work.",
    "1295032": "Hi, \n\nJust wanted to share my idea.  I created a set of unique bssids for each floor; (i.e) for each floor collect bssids which occurs in only that floor. My thinking was that certain bssids are detected in may floors and its difficult to identify the floor based on such commonly occurring bssids. Once such a dictionary is created for each floor in building, for each test path we can find number of bssids matching for each floor data in dictionary. Once such table is created, we can compute data such as count, mean, median of each floor's distribution and select the floor. \n\n\nI have implemented this idea in this [notebook](https://www.kaggle.com/suryajrrafl/wifi-based-floor-mapping/data?scriptVersionId=60842281). 3rd version is my solution, 4th version is the public solution by nigel. The approach I took has 100%accuracy in training set but score's poorly compared to [nigel's solution](https://www.kaggle.com/nigelhenry/simple-99-accurate-floor-model) and found many differences. \n\nAny feedback on the approach is most welcome.",
    "1295201": "aah yes. hadn't thought of majority voting!  let me try that... thanks for the suggestion @mamasinkgs",
    "1295885": "kmldas similarly to you, I agree with the \"99%\" model throughout the public LB (and the 'trick' described in this current thread confirms their correctness) but disagree on 213 points comprising 11 paths in the private LB. Differences are in buildings:\n\n5d2709bb03f801723c32852c (2 paths)\n5d2709d403f801723c32bd39 (3 paths)\n5da138274db8ce0c98bbd3d2 (2 paths)\n5da1382d4db8ce0c98bbe92e (1 path)\n5da138754db8ce0c98bca82f (1 path)\n5da1389e4db8ce0c98bd0547 (1 path)\n5da138b74db8ce0c98bd4774 (1 path)",
    "1296186": "Thanks for sharing @jbomitchell .. good to know we have some differences... this may be critical in the final shakeup....hoping we land up in the right side of it 👍😄",
    "1298536": "In addition to the floor hack, you can also estimate positional bias by adding delta_x, delta_y to your submission. With site  dependent delta (including floor), LB may show which sites are in public or not. Btw, is enough time left to try this?",
    "1298579": "according to this https://www.kaggle.com/c/indoor-location-navigation/discussion/224560 fractional floors are allowed. therefore you can hedge by submitting your (unrounded) mean of all your floor prediction models",
    "1298594": "Thanks @nigelhenry ... i had missed that discussion \nI submitted average of the submissions over 200 were different and yes I get the same score as I think these are all in the hidden set... will know when the results are out which was the better choice... to take average or not!",
    "1301558": "Since the floor penalty uses the absolute difference between the true floor and predicted floor, hedging by predicting the average of the floors is worse in expectation when you are >50% confident of one of the floors."
  },
  "source": "meta"
}