{
  "id": 235813,
  "title": "Hand Labeled Waypoints",
  "url": "/competitions/indoor-location-navigation/discussion/235813",
  "author_name": "",
  "post_date": "2021-05-01T11:04:07.504759100Z",
  "votes": 38,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Our team is using manually generated waypoints for post-processing and I published the dataset.<br>\n<a href=\"https://www.kaggle.com/saitodevel01/indoor-navigation-hand-labeled-waypoints\" target=\"_blank\">Indoor Navigation - Hand Labeled Waypoints</a></p>",
  "messages": [
    {
      "id": "1289778",
      "postDate": "05/01/2021 11:04:07",
      "content": "<p>Our team is using manually generated waypoints for post-processing and I published the dataset.<br>\n<a href=\"https://www.kaggle.com/saitodevel01/indoor-navigation-hand-labeled-waypoints\" target=\"_blank\">Indoor Navigation - Hand Labeled Waypoints</a></p>",
      "rawMarkdown": "Our team is using manually generated waypoints for post-processing and I published the dataset.\n[Indoor Navigation - Hand Labeled Waypoints](https://www.kaggle.com/saitodevel01/indoor-navigation-hand-labeled-waypoints)",
      "votes": null
    },
    {
      "id": "1289839",
      "postDate": "05/01/2021 12:18:01",
      "content": "<p>Thanks for publishing!</p>\n<p>I was a bit worried that this might be close to \"hand labeling\" test data - can anyone comment on that? If you are just looking at the site plans and manually adding waypoints where they seem to be missing, that seems ok (?) (as long as you aren't manually hand labeling the test points?)</p>\n<p>[I'm pretty new to kaggle, so wasn't 100% sure about what would constitute \"hand labeling\" of test data in the rules].  Thanks!</p>",
      "rawMarkdown": "Thanks for publishing!\n\nI was a bit worried that this might be close to \"hand labeling\" test data - can anyone comment on that? If you are just looking at the site plans and manually adding waypoints where they seem to be missing, that seems ok (?) (as long as you aren't manually hand labeling the test points?)\n\n[I'm pretty new to kaggle, so wasn't 100% sure about what would constitute \"hand labeling\" of test data in the rules].  Thanks!",
      "votes": null
    },
    {
      "id": "1289875",
      "postDate": "05/01/2021 12:54:17",
      "content": "<p>I'm worried about that too. I'm wondering if using the data in post-processing is likely to get essentially equivalent results to hand labeling the test waypoints. If this kind of data is allowed to be used, I think the competition will collapse…</p>\n<p>However, I think it will be acceptable, at least if it was generated by some algorithm or machine learning model</p>",
      "rawMarkdown": "I'm worried about that too. I'm wondering if using the data in post-processing is likely to get essentially equivalent results to hand labeling the test waypoints. If this kind of data is allowed to be used, I think the competition will collapse...\n\nHowever, I think it will be acceptable, at least if it was generated by some algorithm or machine learning model",
      "votes": null
    },
    {
      "id": "1289876",
      "postDate": "05/01/2021 12:55:01",
      "content": "<p>Thanks akio for uploading the dataset. <br>\nOur team is also using hand labeled possible (i.e. missing) grid points for the maps and it gave 0.04 improvement in postprocessing. In my understanding, hand labelling on the training dataset is allowed (e.g. <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220522\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220522</a> ). It should be noted that the rainforest competition has the same sentence about the hand-labelling rule:<br>\n<code>Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.</code></p>\n<p><a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/yuanchaoshu\" target=\"_blank\">@yuanchaoshu</a><br>\n1) Is it allowed to use hand labelled possible grid points for maps?<br>\n2) If people use hand labelled possible grid points, does the dataset have to be uploaded?</p>\n<p>For example, the blue points are hand-labelled possible grid points, and the red points are the grid points on the training dataset. <a href=\"https://imgur.com/a/ISU3Dec\" target=\"_blank\">https://imgur.com/a/ISU3Dec</a></p>\n<p>I just want to check my understanding, thanks!</p>\n<p>EDIT: The improvement by hand-labelled possible missing grid point is much bigger now (5/7).</p>",
      "rawMarkdown": "Thanks akio for uploading the dataset. \nOur team is also using hand labeled possible (i.e. missing) grid points for the maps and it gave 0.04 improvement in postprocessing. In my understanding, hand labelling on the training dataset is allowed (e.g. https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220522 ). It should be noted that the rainforest competition has the same sentence about the hand-labelling rule:\n`Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.`\n\n@juliaelliott @addisonhoward @yuanchaoshu\n1) Is it allowed to use hand labelled possible grid points for maps?\n2) If people use hand labelled possible grid points, does the dataset have to be uploaded?\n\nFor example, the blue points are hand-labelled possible grid points, and the red points are the grid points on the training dataset. https://imgur.com/a/ISU3Dec\n\nI just want to check my understanding, thanks!\n\nEDIT: The improvement by hand-labelled possible missing grid point is much bigger now (5/7).",
      "votes": null
    },
    {
      "id": "1289894",
      "postDate": "05/01/2021 13:16:46",
      "content": "<p>This is the visualization of the same sample in the akio's dataset: <a href=\"https://imgur.com/a/U8ulxNy\" target=\"_blank\">https://imgur.com/a/U8ulxNy</a></p>",
      "rawMarkdown": "This is the visualization of the same sample in the akio's dataset: https://imgur.com/a/U8ulxNy",
      "votes": null
    },
    {
      "id": "1289939",
      "postDate": "05/01/2021 13:57:12",
      "content": "<p>I'm also afraid that this can be equivalent to the direct test set labelling, which is against the rule. but personally I think the competition won't collapse anyway, because the prediction below 3 looks very good for me and hard to fix by human. I thought the competition could collapse by hand labelling when the score was around 4, though.</p>",
      "rawMarkdown": "I'm also afraid that this can be equivalent to the direct test set labelling, which is against the rule. but personally I think the competition won't collapse anyway, because the prediction below 3 looks very good for me and hard to fix by human. I thought the competition could collapse by hand labelling when the score was around 4, though.",
      "votes": null
    },
    {
      "id": "1290420",
      "postDate": "05/02/2021 00:11:39",
      "content": "<p>Thank you for sharing <a href=\"https://www.kaggle.com/saitodevel01\" target=\"_blank\">@saitodevel01</a>. As <a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/235733\" target=\"_blank\">I was interested in similar point</a>, it seems very useful.<br>\nWould you elaborate more about how you created this dataset?</p>\n<p>In a generate_grid_points and generate_individual_points functions, there are specific parameters for each site and floor.</p>\n<pre><code>    generate_grid_points(\n    site  = '5d2709d403f801723c32bd39',\n    floor = 0,\n    top_left     = np.array((86.4, 135.0)),\n    top_right    = np.array((93.4, 133.7)),\n    bottom_left  = np.array((83.9, 117.3)),\n    bottom_right = np.array((88.7, 116.9)),\n    N_h = 3,\n    N_v = 10,\n    )\n</code></pre>\n<pre><code>    generate_individual_points(\n        site  = '5da1382d4db8ce0c98bbe92e',\n        floor = 1,\n        points = [\n            (126, 160),\n            (114.5, 162),\n            (114.5, 156),\n            (119.5, 156),\n            (119.5, 147.5),\n            (107, 157),\n            (107.5, 145.5),\n            (113.5, 145.5),\n        ])\n</code></pre>\n<p>Where these values come from?<br>\nDid you define them based solely on train waypoints and floor maps (and not based on test data,  LB score, or other external data)? </p>",
      "rawMarkdown": "Thank you for sharing @saitodevel01. As [I was interested in similar point](https://www.kaggle.com/c/indoor-location-navigation/discussion/235733), it seems very useful.\nWould you elaborate more about how you created this dataset?\n\nIn a generate_grid_points and generate_individual_points functions, there are specific parameters for each site and floor.\n```\n    generate_grid_points(\n    site  = '5d2709d403f801723c32bd39',\n    floor = 0,\n    top_left     = np.array((86.4, 135.0)),\n    top_right    = np.array((93.4, 133.7)),\n    bottom_left  = np.array((83.9, 117.3)),\n    bottom_right = np.array((88.7, 116.9)),\n    N_h = 3,\n    N_v = 10,\n    )\n```\n ```\n    generate_individual_points(\n        site  = '5da1382d4db8ce0c98bbe92e',\n        floor = 1,\n        points = [\n            (126, 160),\n            (114.5, 162),\n            (114.5, 156),\n            (119.5, 156),\n            (119.5, 147.5),\n            (107, 157),\n            (107.5, 145.5),\n            (113.5, 145.5),\n        ])\n```\n\nWhere these values come from?\nDid you define them based solely on train waypoints and floor maps (and not based on test data,  LB score, or other external data)?",
      "votes": null
    },
    {
      "id": "1290520",
      "postDate": "05/02/2021 04:49:47",
      "content": "<p>I manually added possible missing grid points around my predictions using floor plan. I'm also afraid that this can be equivalent to the direct test set labelling.</p>",
      "rawMarkdown": "I manually added possible missing grid points around my predictions using floor plan. I'm also afraid that this can be equivalent to the direct test set labelling.",
      "votes": null
    },
    {
      "id": "1290541",
      "postDate": "05/02/2021 05:25:09",
      "content": "<p>Thank you for clarifying your method.<br>\nAdding new waypoints based on test predictions seems to be a bit dangerous if it is incorporated with  <a href=\"https://www.kaggle.com/robikscube/indoor-navigation-snap-to-grid-post-processing\" target=\"_blank\">snap-to-grid</a>. It is practically similar to manually editing the test predictions. <br>\nI also think that adding new waypoints is not hand labeling of train set, as they are not the labels of train data actually. I am not sure it should be regarded as hand labeling of test set, but at least, new waypoints are the candidates of test labels that are not in train ones.</p>\n<p>As I am a novice, I am not sure about it. I am happy when anyone correct me if I am wrong. <br>\nI never meant to blame any teams using hand labeled waypoints. </p>",
      "rawMarkdown": "Thank you for clarifying your method.\nAdding new waypoints based on test predictions seems to be a bit dangerous if it is incorporated with  [snap-to-grid](https://www.kaggle.com/robikscube/indoor-navigation-snap-to-grid-post-processing). It is practically similar to manually editing the test predictions. \nI also think that adding new waypoints is not hand labeling of train set, as they are not the labels of train data actually. I am not sure it should be regarded as hand labeling of test set, but at least, new waypoints are the candidates of test labels that are not in train ones.\n\nAs I am a novice, I am not sure about it. I am happy when anyone correct me if I am wrong. \nI never meant to blame any teams using hand labeled waypoints.",
      "votes": null
    },
    {
      "id": "1290882",
      "postDate": "05/02/2021 14:00:14",
      "content": "<p>As <a href=\"https://www.kaggle.com/tomooinubushi\" target=\"_blank\">@tomooinubushi</a> said, adding waypoints based on test predictions seems to be somewhat gray, as it can be close to the hand labelling of the test set. However, as <a href=\"https://www.kaggle.com/chris62\" target=\"_blank\">@chris62</a> says, I think manually adding missing waypoints by looking at the map is ok, because it doesn't use the information of the test set, so not the hand labelling of the test set. My teammate, Reza, who made extra_waypoints.csv, said he generated missing way_point for the coverage hole. So I'm going to use our extra_waypoints.csv if <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/yuanchaoshu\" target=\"_blank\">@yuanchaoshu</a> doesn't prohibit the use of possible missing grid points. Maybe we will use the 2 versions (with/without extra_waypoints.csv) for the final submissions for the safety, though.</p>",
      "rawMarkdown": "As @tomooinubushi said, adding waypoints based on test predictions seems to be somewhat gray, as it can be close to the hand labelling of the test set. However, as @chris62 says, I think manually adding missing waypoints by looking at the map is ok, because it doesn't use the information of the test set, so not the hand labelling of the test set. My teammate, Reza, who made extra_waypoints.csv, said he generated missing way_point for the coverage hole. So I'm going to use our extra_waypoints.csv if @juliaelliott @addisonhoward @yuanchaoshu doesn't prohibit the use of possible missing grid points. Maybe we will use the 2 versions (with/without extra_waypoints.csv) for the final submissions for the safety, though.",
      "votes": null
    },
    {
      "id": "1295892",
      "postDate": "05/06/2021 19:26:31",
      "content": "<p>Does anyone else have more thoughts about this, or has there been any official word from Kaggle about it?</p>",
      "rawMarkdown": "Does anyone else have more thoughts about this, or has there been any official word from Kaggle about it?",
      "votes": null
    },
    {
      "id": "1296827",
      "postDate": "05/07/2021 14:18:36",
      "content": "<p>It is prohibited by the rules to hand label test data (this includes the public and private test set). However, it is permitted to relabel or handlabel the training set.</p>",
      "rawMarkdown": "It is prohibited by the rules to hand label test data (this includes the public and private test set). However, it is permitted to relabel or handlabel the training set.",
      "votes": null
    },
    {
      "id": "1296978",
      "postDate": "05/07/2021 16:06:17",
      "content": "<p><a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> Thanks you for answering. I want to check one.  Adding possible missing grid points by hand for post processing is permitted? (like <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> show <a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/235813#1289894\" target=\"_blank\">here</a>) </p>",
      "rawMarkdown": "juliaelliott Thanks you for answering. I want to check one.  Adding possible missing grid points by hand for post processing is permitted? (like @mamasinkgs show [here](https://www.kaggle.com/c/indoor-location-navigation/discussion/235813#1289894))",
      "votes": null
    },
    {
      "id": "1296983",
      "postDate": "05/07/2021 16:14:46",
      "content": "<p>I\"m sorry, the link above didn't work.<br>\nWhat I mentioned is following thread.</p>\n<blockquote>\n  <p>Thanks akio for uploading the dataset. <br>\n  Our team is also using hand labeled possible (i.e. missing) grid points for the maps and it gave 0.04 improvement in postprocessing. In my understanding, hand labelling on the training dataset is allowed (e.g. <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220522\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220522</a> ). It should be noted that the rainforest competition has the same sentence about the hand-labelling rule:<br>\n  <code>Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.</code></p>\n  <p><a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/yuanchaoshu\" target=\"_blank\">@yuanchaoshu</a><br>\n  1) Is it allowed to use hand labelled possible grid points for maps?<br>\n  2) If people use hand labelled possible grid points, does the dataset have to be uploaded?</p>\n  <p>For example, the blue points are hand-labelled possible grid points, and the red points are the grid points on the training dataset. <a href=\"https://imgur.com/a/ISU3Dec\" target=\"_blank\">https://imgur.com/a/ISU3Dec</a></p>\n  <p>I just want to check my understanding, thanks!</p>\n  <p>EDIT: The improvement by hand-labelled possible missing grid point is much bigger now (5/7).</p>\n</blockquote>",
      "rawMarkdown": "I\"m sorry, the link above didn't work.\nWhat I mentioned is following thread.\n\n> Thanks akio for uploading the dataset. \n> Our team is also using hand labeled possible (i.e. missing) grid points for the maps and it gave 0.04 improvement in postprocessing. In my understanding, hand labelling on the training dataset is allowed (e.g. https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220522 ). It should be noted that the rainforest competition has the same sentence about the hand-labelling rule:\n> `Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.`\n> \n> @juliaelliott @addisonhoward @yuanchaoshu\n> 1) Is it allowed to use hand labelled possible grid points for maps?\n> 2) If people use hand labelled possible grid points, does the dataset have to be uploaded?\n> \n> For example, the blue points are hand-labelled possible grid points, and the red points are the grid points on the training dataset. https://imgur.com/a/ISU3Dec\n> \n> I just want to check my understanding, thanks!\n> \n> EDIT: The improvement by hand-labelled possible missing grid point is much bigger now (5/7).",
      "votes": null
    },
    {
      "id": "1297203",
      "postDate": "05/07/2021 19:40:29",
      "content": "<p>Yes, I'm sure all the teams are looking forward to the kaggle team and host ( <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/yuanchaoshu\" target=\"_blank\">@yuanchaoshu</a>)'s clarification of the use of hand-labelled possible missing grid points for the postprocessing.</p>",
      "rawMarkdown": "Yes, I'm sure all the teams are looking forward to the kaggle team and host ( @juliaelliott @addisonhoward @yuanchaoshu)'s clarification of the use of hand-labelled possible missing grid points for the postprocessing.",
      "votes": null
    },
    {
      "id": "1297306",
      "postDate": "05/07/2021 22:40:12",
      "content": "<p>The guiding general principle is this: \"Would this method still be applicable to a completely new, unseen test set?\" If the answer is YES, then it is allowable. If the answer is, NO then it is not allowable.</p>\n<p>So, in this case, <em>hand</em> labeling of grid points that would be used for Test traces would not be acceptable. An automated process to fill in grid points that would work for new locations would be acceptable.</p>",
      "rawMarkdown": "The guiding general principle is this: \"Would this method still be applicable to a completely new, unseen test set?\" If the answer is YES, then it is allowable. If the answer is, NO then it is not allowable.\n\nSo, in this case, *hand* labeling of grid points that would be used for Test traces would not be acceptable. An automated process to fill in grid points that would work for new locations would be acceptable.",
      "votes": null
    },
    {
      "id": "1297320",
      "postDate": "05/07/2021 23:10:07",
      "content": "<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> Thank you for the clarification! Then, our team won't choose the sub that uses the hand-labelled missing grid points for the final submission. </p>",
      "rawMarkdown": "inversion Thank you for the clarification! Then, our team won't choose the sub that uses the hand-labelled missing grid points for the final submission.",
      "votes": null
    },
    {
      "id": "1297321",
      "postDate": "05/07/2021 23:13:12",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> for the clarification. We won't use submissions with hand labeled missing grid points either.</p>",
      "rawMarkdown": "Thank you @inversion for the clarification. We won't use submissions with hand labeled missing grid points either.",
      "votes": null
    },
    {
      "id": "1297325",
      "postDate": "05/07/2021 23:21:07",
      "content": "<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> Thank you very much. I understand it. Our team will no longer use hand labeling waypoints.</p>",
      "rawMarkdown": "inversion Thank you very much. I understand it. Our team will no longer use hand labeling waypoints.",
      "votes": null
    },
    {
      "id": "1297392",
      "postDate": "05/08/2021 01:23:02",
      "content": "<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a><br>\nThank you.<br>\nI think that in Saito's dataset, grid points generated by \"generate_grid_points\" is still OK because they are generated automatically, but points generated by \"generate_individual_points\" are NG because they are manually.<br>\nor am I wrong and both are NG?</p>",
      "rawMarkdown": "inversion\nThank you.\nI think that in Saito's dataset, grid points generated by \"generate_grid_points\" is still OK because they are generated automatically, but points generated by \"generate_individual_points\" are NG because they are manually.\nor am I wrong and both are NG?",
      "votes": null
    },
    {
      "id": "1297491",
      "postDate": "05/08/2021 04:48:05",
      "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a>.</p>",
      "rawMarkdown": "Thank you very much @inversion.",
      "votes": null
    },
    {
      "id": "1297492",
      "postDate": "05/08/2021 04:48:27",
      "content": "<p>Our team has developed an automated process for generating additional grid points and generate_grid_points.py is no longer used.</p>",
      "rawMarkdown": "Our team has developed an automated process for generating additional grid points and generate_grid_points.py is no longer used.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1289839,
      "author_name": "chris62",
      "author_url": "",
      "post_date": "05/01/2021 12:18:01",
      "content": "<p>Thanks for publishing!</p>\n<p>I was a bit worried that this might be close to \"hand labeling\" test data - can anyone comment on that? If you are just looking at the site plans and manually adding waypoints where they seem to be missing, that seems ok (?) (as long as you aren't manually hand labeling the test points?)</p>\n<p>[I'm pretty new to kaggle, so wasn't 100% sure about what would constitute \"hand labeling\" of test data in the rules].  Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1290882,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "05/02/2021 14:00:14",
          "content": "<p>As <a href=\"https://www.kaggle.com/tomooinubushi\" target=\"_blank\">@tomooinubushi</a> said, adding waypoints based on test predictions seems to be somewhat gray, as it can be close to the hand labelling of the test set. However, as <a href=\"https://www.kaggle.com/chris62\" target=\"_blank\">@chris62</a> says, I think manually adding missing waypoints by looking at the map is ok, because it doesn't use the information of the test set, so not the hand labelling of the test set. My teammate, Reza, who made extra_waypoints.csv, said he generated missing way_point for the coverage hole. So I'm going to use our extra_waypoints.csv if <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/yuanchaoshu\" target=\"_blank\">@yuanchaoshu</a> doesn't prohibit the use of possible missing grid points. Maybe we will use the 2 versions (with/without extra_waypoints.csv) for the final submissions for the safety, though.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1289875,
      "author_name": "yamsam",
      "author_url": "",
      "post_date": "05/01/2021 12:54:17",
      "content": "<p>I'm worried about that too. I'm wondering if using the data in post-processing is likely to get essentially equivalent results to hand labeling the test waypoints. If this kind of data is allowed to be used, I think the competition will collapse…</p>\n<p>However, I think it will be acceptable, at least if it was generated by some algorithm or machine learning model</p>",
      "votes": null,
      "replies": [
        {
          "id": 1289939,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "05/01/2021 13:57:12",
          "content": "<p>I'm also afraid that this can be equivalent to the direct test set labelling, which is against the rule. but personally I think the competition won't collapse anyway, because the prediction below 3 looks very good for me and hard to fix by human. I thought the competition could collapse by hand labelling when the score was around 4, though.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1289876,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "05/01/2021 12:55:01",
      "content": "<p>Thanks akio for uploading the dataset. <br>\nOur team is also using hand labeled possible (i.e. missing) grid points for the maps and it gave 0.04 improvement in postprocessing. In my understanding, hand labelling on the training dataset is allowed (e.g. <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220522\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220522</a> ). It should be noted that the rainforest competition has the same sentence about the hand-labelling rule:<br>\n<code>Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.</code></p>\n<p><a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/yuanchaoshu\" target=\"_blank\">@yuanchaoshu</a><br>\n1) Is it allowed to use hand labelled possible grid points for maps?<br>\n2) If people use hand labelled possible grid points, does the dataset have to be uploaded?</p>\n<p>For example, the blue points are hand-labelled possible grid points, and the red points are the grid points on the training dataset. <a href=\"https://imgur.com/a/ISU3Dec\" target=\"_blank\">https://imgur.com/a/ISU3Dec</a></p>\n<p>I just want to check my understanding, thanks!</p>\n<p>EDIT: The improvement by hand-labelled possible missing grid point is much bigger now (5/7).</p>",
      "votes": null,
      "replies": [
        {
          "id": 1289894,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "05/01/2021 13:16:46",
          "content": "<p>This is the visualization of the same sample in the akio's dataset: <a href=\"https://imgur.com/a/U8ulxNy\" target=\"_blank\">https://imgur.com/a/U8ulxNy</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1290420,
      "author_name": "tomooinubushi",
      "author_url": "",
      "post_date": "05/02/2021 00:11:39",
      "content": "<p>Thank you for sharing <a href=\"https://www.kaggle.com/saitodevel01\" target=\"_blank\">@saitodevel01</a>. As <a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/235733\" target=\"_blank\">I was interested in similar point</a>, it seems very useful.<br>\nWould you elaborate more about how you created this dataset?</p>\n<p>In a generate_grid_points and generate_individual_points functions, there are specific parameters for each site and floor.</p>\n<pre><code>    generate_grid_points(\n    site  = '5d2709d403f801723c32bd39',\n    floor = 0,\n    top_left     = np.array((86.4, 135.0)),\n    top_right    = np.array((93.4, 133.7)),\n    bottom_left  = np.array((83.9, 117.3)),\n    bottom_right = np.array((88.7, 116.9)),\n    N_h = 3,\n    N_v = 10,\n    )\n</code></pre>\n<pre><code>    generate_individual_points(\n        site  = '5da1382d4db8ce0c98bbe92e',\n        floor = 1,\n        points = [\n            (126, 160),\n            (114.5, 162),\n            (114.5, 156),\n            (119.5, 156),\n            (119.5, 147.5),\n            (107, 157),\n            (107.5, 145.5),\n            (113.5, 145.5),\n        ])\n</code></pre>\n<p>Where these values come from?<br>\nDid you define them based solely on train waypoints and floor maps (and not based on test data,  LB score, or other external data)? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1290520,
      "author_name": "saitodevel01",
      "author_url": "",
      "post_date": "05/02/2021 04:49:47",
      "content": "<p>I manually added possible missing grid points around my predictions using floor plan. I'm also afraid that this can be equivalent to the direct test set labelling.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1290541,
          "author_name": "tomooinubushi",
          "author_url": "",
          "post_date": "05/02/2021 05:25:09",
          "content": "<p>Thank you for clarifying your method.<br>\nAdding new waypoints based on test predictions seems to be a bit dangerous if it is incorporated with  <a href=\"https://www.kaggle.com/robikscube/indoor-navigation-snap-to-grid-post-processing\" target=\"_blank\">snap-to-grid</a>. It is practically similar to manually editing the test predictions. <br>\nI also think that adding new waypoints is not hand labeling of train set, as they are not the labels of train data actually. I am not sure it should be regarded as hand labeling of test set, but at least, new waypoints are the candidates of test labels that are not in train ones.</p>\n<p>As I am a novice, I am not sure about it. I am happy when anyone correct me if I am wrong. <br>\nI never meant to blame any teams using hand labeled waypoints. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1295892,
      "author_name": "chris62",
      "author_url": "",
      "post_date": "05/06/2021 19:26:31",
      "content": "<p>Does anyone else have more thoughts about this, or has there been any official word from Kaggle about it?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1296827,
      "author_name": "juliaelliott",
      "author_url": "",
      "post_date": "05/07/2021 14:18:36",
      "content": "<p>It is prohibited by the rules to hand label test data (this includes the public and private test set). However, it is permitted to relabel or handlabel the training set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1296978,
          "author_name": "kuto0633",
          "author_url": "",
          "post_date": "05/07/2021 16:06:17",
          "content": "<p><a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> Thanks you for answering. I want to check one.  Adding possible missing grid points by hand for post processing is permitted? (like <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> show <a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/235813#1289894\" target=\"_blank\">here</a>) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1296983,
          "author_name": "kuto0633",
          "author_url": "",
          "post_date": "05/07/2021 16:14:46",
          "content": "<p>I\"m sorry, the link above didn't work.<br>\nWhat I mentioned is following thread.</p>\n<blockquote>\n  <p>Thanks akio for uploading the dataset. <br>\n  Our team is also using hand labeled possible (i.e. missing) grid points for the maps and it gave 0.04 improvement in postprocessing. In my understanding, hand labelling on the training dataset is allowed (e.g. <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220522\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220522</a> ). It should be noted that the rainforest competition has the same sentence about the hand-labelling rule:<br>\n  <code>Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.</code></p>\n  <p><a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/yuanchaoshu\" target=\"_blank\">@yuanchaoshu</a><br>\n  1) Is it allowed to use hand labelled possible grid points for maps?<br>\n  2) If people use hand labelled possible grid points, does the dataset have to be uploaded?</p>\n  <p>For example, the blue points are hand-labelled possible grid points, and the red points are the grid points on the training dataset. <a href=\"https://imgur.com/a/ISU3Dec\" target=\"_blank\">https://imgur.com/a/ISU3Dec</a></p>\n  <p>I just want to check my understanding, thanks!</p>\n  <p>EDIT: The improvement by hand-labelled possible missing grid point is much bigger now (5/7).</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1297203,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "05/07/2021 19:40:29",
          "content": "<p>Yes, I'm sure all the teams are looking forward to the kaggle team and host ( <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/yuanchaoshu\" target=\"_blank\">@yuanchaoshu</a>)'s clarification of the use of hand-labelled possible missing grid points for the postprocessing.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1297306,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "05/07/2021 22:40:12",
          "content": "<p>The guiding general principle is this: \"Would this method still be applicable to a completely new, unseen test set?\" If the answer is YES, then it is allowable. If the answer is, NO then it is not allowable.</p>\n<p>So, in this case, <em>hand</em> labeling of grid points that would be used for Test traces would not be acceptable. An automated process to fill in grid points that would work for new locations would be acceptable.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1297320,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "05/07/2021 23:10:07",
          "content": "<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> Thank you for the clarification! Then, our team won't choose the sub that uses the hand-labelled missing grid points for the final submission. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1297321,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "05/07/2021 23:13:12",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> for the clarification. We won't use submissions with hand labeled missing grid points either.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1297325,
          "author_name": "kuto0633",
          "author_url": "",
          "post_date": "05/07/2021 23:21:07",
          "content": "<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> Thank you very much. I understand it. Our team will no longer use hand labeling waypoints.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1297392,
          "author_name": "iwatatakuya",
          "author_url": "",
          "post_date": "05/08/2021 01:23:02",
          "content": "<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a><br>\nThank you.<br>\nI think that in Saito's dataset, grid points generated by \"generate_grid_points\" is still OK because they are generated automatically, but points generated by \"generate_individual_points\" are NG because they are manually.<br>\nor am I wrong and both are NG?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1297491,
          "author_name": "saitodevel01",
          "author_url": "",
          "post_date": "05/08/2021 04:48:05",
          "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1297492,
          "author_name": "saitodevel01",
          "author_url": "",
          "post_date": "05/08/2021 04:48:27",
          "content": "<p>Our team has developed an automated process for generating additional grid points and generate_grid_points.py is no longer used.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1289778": "Our team is using manually generated waypoints for post-processing and I published the dataset.\n[Indoor Navigation - Hand Labeled Waypoints](https://www.kaggle.com/saitodevel01/indoor-navigation-hand-labeled-waypoints)",
    "1289839": "Thanks for publishing!\n\nI was a bit worried that this might be close to \"hand labeling\" test data - can anyone comment on that? If you are just looking at the site plans and manually adding waypoints where they seem to be missing, that seems ok (?) (as long as you aren't manually hand labeling the test points?)\n\n[I'm pretty new to kaggle, so wasn't 100% sure about what would constitute \"hand labeling\" of test data in the rules].  Thanks!",
    "1289875": "I'm worried about that too. I'm wondering if using the data in post-processing is likely to get essentially equivalent results to hand labeling the test waypoints. If this kind of data is allowed to be used, I think the competition will collapse...\n\nHowever, I think it will be acceptable, at least if it was generated by some algorithm or machine learning model",
    "1289876": "Thanks akio for uploading the dataset. \nOur team is also using hand labeled possible (i.e. missing) grid points for the maps and it gave 0.04 improvement in postprocessing. In my understanding, hand labelling on the training dataset is allowed (e.g. https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220522 ). It should be noted that the rainforest competition has the same sentence about the hand-labelling rule:\n`Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.`\n\n@juliaelliott @addisonhoward @yuanchaoshu\n1) Is it allowed to use hand labelled possible grid points for maps?\n2) If people use hand labelled possible grid points, does the dataset have to be uploaded?\n\nFor example, the blue points are hand-labelled possible grid points, and the red points are the grid points on the training dataset. https://imgur.com/a/ISU3Dec\n\nI just want to check my understanding, thanks!\n\nEDIT: The improvement by hand-labelled possible missing grid point is much bigger now (5/7).",
    "1289894": "This is the visualization of the same sample in the akio's dataset: https://imgur.com/a/U8ulxNy",
    "1289939": "I'm also afraid that this can be equivalent to the direct test set labelling, which is against the rule. but personally I think the competition won't collapse anyway, because the prediction below 3 looks very good for me and hard to fix by human. I thought the competition could collapse by hand labelling when the score was around 4, though.",
    "1290420": "Thank you for sharing @saitodevel01. As [I was interested in similar point](https://www.kaggle.com/c/indoor-location-navigation/discussion/235733), it seems very useful.\nWould you elaborate more about how you created this dataset?\n\nIn a generate_grid_points and generate_individual_points functions, there are specific parameters for each site and floor.\n```\n    generate_grid_points(\n    site  = '5d2709d403f801723c32bd39',\n    floor = 0,\n    top_left     = np.array((86.4, 135.0)),\n    top_right    = np.array((93.4, 133.7)),\n    bottom_left  = np.array((83.9, 117.3)),\n    bottom_right = np.array((88.7, 116.9)),\n    N_h = 3,\n    N_v = 10,\n    )\n```\n ```\n    generate_individual_points(\n        site  = '5da1382d4db8ce0c98bbe92e',\n        floor = 1,\n        points = [\n            (126, 160),\n            (114.5, 162),\n            (114.5, 156),\n            (119.5, 156),\n            (119.5, 147.5),\n            (107, 157),\n            (107.5, 145.5),\n            (113.5, 145.5),\n        ])\n```\n\nWhere these values come from?\nDid you define them based solely on train waypoints and floor maps (and not based on test data,  LB score, or other external data)?",
    "1290520": "I manually added possible missing grid points around my predictions using floor plan. I'm also afraid that this can be equivalent to the direct test set labelling.",
    "1290541": "Thank you for clarifying your method.\nAdding new waypoints based on test predictions seems to be a bit dangerous if it is incorporated with  [snap-to-grid](https://www.kaggle.com/robikscube/indoor-navigation-snap-to-grid-post-processing). It is practically similar to manually editing the test predictions. \nI also think that adding new waypoints is not hand labeling of train set, as they are not the labels of train data actually. I am not sure it should be regarded as hand labeling of test set, but at least, new waypoints are the candidates of test labels that are not in train ones.\n\nAs I am a novice, I am not sure about it. I am happy when anyone correct me if I am wrong. \nI never meant to blame any teams using hand labeled waypoints.",
    "1290882": "As @tomooinubushi said, adding waypoints based on test predictions seems to be somewhat gray, as it can be close to the hand labelling of the test set. However, as @chris62 says, I think manually adding missing waypoints by looking at the map is ok, because it doesn't use the information of the test set, so not the hand labelling of the test set. My teammate, Reza, who made extra_waypoints.csv, said he generated missing way_point for the coverage hole. So I'm going to use our extra_waypoints.csv if @juliaelliott @addisonhoward @yuanchaoshu doesn't prohibit the use of possible missing grid points. Maybe we will use the 2 versions (with/without extra_waypoints.csv) for the final submissions for the safety, though.",
    "1295892": "Does anyone else have more thoughts about this, or has there been any official word from Kaggle about it?",
    "1296827": "It is prohibited by the rules to hand label test data (this includes the public and private test set). However, it is permitted to relabel or handlabel the training set.",
    "1296978": "juliaelliott Thanks you for answering. I want to check one.  Adding possible missing grid points by hand for post processing is permitted? (like @mamasinkgs show [here](https://www.kaggle.com/c/indoor-location-navigation/discussion/235813#1289894))",
    "1296983": "I\"m sorry, the link above didn't work.\nWhat I mentioned is following thread.\n\n> Thanks akio for uploading the dataset. \n> Our team is also using hand labeled possible (i.e. missing) grid points for the maps and it gave 0.04 improvement in postprocessing. In my understanding, hand labelling on the training dataset is allowed (e.g. https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220522 ). It should be noted that the rainforest competition has the same sentence about the hand-labelling rule:\n> `Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.`\n> \n> @juliaelliott @addisonhoward @yuanchaoshu\n> 1) Is it allowed to use hand labelled possible grid points for maps?\n> 2) If people use hand labelled possible grid points, does the dataset have to be uploaded?\n> \n> For example, the blue points are hand-labelled possible grid points, and the red points are the grid points on the training dataset. https://imgur.com/a/ISU3Dec\n> \n> I just want to check my understanding, thanks!\n> \n> EDIT: The improvement by hand-labelled possible missing grid point is much bigger now (5/7).",
    "1297203": "Yes, I'm sure all the teams are looking forward to the kaggle team and host ( @juliaelliott @addisonhoward @yuanchaoshu)'s clarification of the use of hand-labelled possible missing grid points for the postprocessing.",
    "1297306": "The guiding general principle is this: \"Would this method still be applicable to a completely new, unseen test set?\" If the answer is YES, then it is allowable. If the answer is, NO then it is not allowable.\n\nSo, in this case, *hand* labeling of grid points that would be used for Test traces would not be acceptable. An automated process to fill in grid points that would work for new locations would be acceptable.",
    "1297320": "inversion Thank you for the clarification! Then, our team won't choose the sub that uses the hand-labelled missing grid points for the final submission.",
    "1297321": "Thank you @inversion for the clarification. We won't use submissions with hand labeled missing grid points either.",
    "1297325": "inversion Thank you very much. I understand it. Our team will no longer use hand labeling waypoints.",
    "1297392": "inversion\nThank you.\nI think that in Saito's dataset, grid points generated by \"generate_grid_points\" is still OK because they are generated automatically, but points generated by \"generate_individual_points\" are NG because they are manually.\nor am I wrong and both are NG?",
    "1297491": "Thank you very much @inversion.",
    "1297492": "Our team has developed an automated process for generating additional grid points and generate_grid_points.py is no longer used."
  },
  "source": "meta"
}