{
  "id": 239025,
  "title": "LB score before snap-to-grid",
  "url": "/competitions/indoor-location-navigation/discussion/239025",
  "author_name": "",
  "post_date": "2021-05-14T10:48:24.435869400Z",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I am curious how much improvement can be achieved with proper snap-to-grid. My submission LB scores before / after snap-to-grid (which is my final post-processing step) are: <strong>4.457 / 3.806</strong>. You are welcome to share your results if you want.</p>\n<p>I am still investigating why snapping points that are predicted to be in the hallways to the training waypoints results in <strong>better performance</strong> than leaving them not post processed. Maybe training waypoints appear also in testing waypoints? Any thoughts on that? :)</p>",
  "messages": [
    {
      "id": "1307253",
      "postDate": "05/14/2021 10:48:24",
      "content": "<p>I am curious how much improvement can be achieved with proper snap-to-grid. My submission LB scores before / after snap-to-grid (which is my final post-processing step) are: <strong>4.457 / 3.806</strong>. You are welcome to share your results if you want.</p>\n<p>I am still investigating why snapping points that are predicted to be in the hallways to the training waypoints results in <strong>better performance</strong> than leaving them not post processed. Maybe training waypoints appear also in testing waypoints? Any thoughts on that? :)</p>",
      "rawMarkdown": "I am curious how much improvement can be achieved with proper snap-to-grid. My submission LB scores before / after snap-to-grid (which is my final post-processing step) are: **4.457 / 3.806**. You are welcome to share your results if you want.\n\nI am still investigating why snapping points that are predicted to be in the hallways to the training waypoints results in **better performance** than leaving them not post processed. Maybe training waypoints appear also in testing waypoints? Any thoughts on that? :)",
      "votes": null
    },
    {
      "id": "1307257",
      "postDate": "05/14/2021 10:52:32",
      "content": "<p>Waypoints definitely overlap between the train and test sets.</p>",
      "rawMarkdown": "Waypoints definitely overlap between the train and test sets.",
      "votes": null
    },
    {
      "id": "1307390",
      "postDate": "05/14/2021 12:22:44",
      "content": "<p>Like <a href=\"https://www.kaggle.com/tvdwiele\" target=\"_blank\">@tvdwiele</a> mentioned, train and test waypoints overlap quite a bit - my ballpark estimate is something like 80% - 90% of the test waypoints are also in the training data, so that's why it's advantageous to snap to training points.</p>\n<p>Snapping used to get me ~0.5m or so, but as my model and other post-processing has gotten better, snapping has become less and less important (though I still do a \"modified snap\" as a final step)</p>",
      "rawMarkdown": "Like @tvdwiele mentioned, train and test waypoints overlap quite a bit - my ballpark estimate is something like 80% - 90% of the test waypoints are also in the training data, so that's why it's advantageous to snap to training points.\n\nSnapping used to get me ~0.5m or so, but as my model and other post-processing has gotten better, snapping has become less and less important (though I still do a \"modified snap\" as a final step)",
      "votes": null
    },
    {
      "id": "1307934",
      "postDate": "05/14/2021 19:31:36",
      "content": "<p>The short answer is: Snap to Grid 4.235 -&gt; 4.050</p>\n<p>A longer answer is that I'm doing an iterative looping scheme involving different postprocessing steps (cost minimization, leakage, snap-to-grid, but not always in that order), and the score moves up and down with each step as I go around the loops. Once the full loop stops improving the score, then I give up and submit the best score I found. In general, snapping to the training grid gains me about 0.2 and is better (for me) than snapping to the full generated grid.</p>",
      "rawMarkdown": "The short answer is: Snap to Grid 4.235 -> 4.050\n\nA longer answer is that I'm doing an iterative looping scheme involving different postprocessing steps (cost minimization, leakage, snap-to-grid, but not always in that order), and the score moves up and down with each step as I go around the loops. Once the full loop stops improving the score, then I give up and submit the best score I found. In general, snapping to the training grid gains me about 0.2 and is better (for me) than snapping to the full generated grid.",
      "votes": null
    },
    {
      "id": "1310366",
      "postDate": "05/16/2021 15:55:23",
      "content": "<blockquote>\n  <p>I am still investigating why snapping points that are predicted to be in the hallways to the training waypoints results in better performance than leaving them not post processed.</p>\n</blockquote>\n<p>This competition is not about predicting exact x,y coordinates- it is about predicting which pre-defined location was traveled to. As already mentioned, in 80%-90% of the test set we already know these pre-defined locations.</p>\n<blockquote>\n  <p>Maybe training waypoints appear also in testing waypoints?</p>\n</blockquote>\n<p>Exactly, This is part of why snapping works so well. I'm surprised at how strong your score is without knowing this already. Great work on your current LB score!</p>\n<p>Wish I had more free time to have worked on this competition. Lots of fun ideas I would've liked to try.</p>",
      "rawMarkdown": "> I am still investigating why snapping points that are predicted to be in the hallways to the training waypoints results in better performance than leaving them not post processed.\n\nThis competition is not about predicting exact x,y coordinates- it is about predicting which pre-defined location was traveled to. As already mentioned, in 80%-90% of the test set we already know these pre-defined locations.\n\n> Maybe training waypoints appear also in testing waypoints?\n\nExactly, This is part of why snapping works so well. I'm surprised at how strong your score is without knowing this already. Great work on your current LB score!\n\nWish I had more free time to have worked on this competition. Lots of fun ideas I would've liked to try.",
      "votes": null
    },
    {
      "id": "1310400",
      "postDate": "05/16/2021 16:23:46",
      "content": "<p>Thank you for the reply :) This is the most important point that I was missing. </p>",
      "rawMarkdown": "Thank you for the reply :) This is the most important point that I was missing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1307257,
      "author_name": "tvdwiele",
      "author_url": "",
      "post_date": "05/14/2021 10:52:32",
      "content": "<p>Waypoints definitely overlap between the train and test sets.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1307390,
      "author_name": "chris62",
      "author_url": "",
      "post_date": "05/14/2021 12:22:44",
      "content": "<p>Like <a href=\"https://www.kaggle.com/tvdwiele\" target=\"_blank\">@tvdwiele</a> mentioned, train and test waypoints overlap quite a bit - my ballpark estimate is something like 80% - 90% of the test waypoints are also in the training data, so that's why it's advantageous to snap to training points.</p>\n<p>Snapping used to get me ~0.5m or so, but as my model and other post-processing has gotten better, snapping has become less and less important (though I still do a \"modified snap\" as a final step)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1307934,
      "author_name": "jbomitchell",
      "author_url": "",
      "post_date": "05/14/2021 19:31:36",
      "content": "<p>The short answer is: Snap to Grid 4.235 -&gt; 4.050</p>\n<p>A longer answer is that I'm doing an iterative looping scheme involving different postprocessing steps (cost minimization, leakage, snap-to-grid, but not always in that order), and the score moves up and down with each step as I go around the loops. Once the full loop stops improving the score, then I give up and submit the best score I found. In general, snapping to the training grid gains me about 0.2 and is better (for me) than snapping to the full generated grid.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1310366,
      "author_name": "robikscube",
      "author_url": "",
      "post_date": "05/16/2021 15:55:23",
      "content": "<blockquote>\n  <p>I am still investigating why snapping points that are predicted to be in the hallways to the training waypoints results in better performance than leaving them not post processed.</p>\n</blockquote>\n<p>This competition is not about predicting exact x,y coordinates- it is about predicting which pre-defined location was traveled to. As already mentioned, in 80%-90% of the test set we already know these pre-defined locations.</p>\n<blockquote>\n  <p>Maybe training waypoints appear also in testing waypoints?</p>\n</blockquote>\n<p>Exactly, This is part of why snapping works so well. I'm surprised at how strong your score is without knowing this already. Great work on your current LB score!</p>\n<p>Wish I had more free time to have worked on this competition. Lots of fun ideas I would've liked to try.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1310400,
          "author_name": "olaf2000",
          "author_url": "",
          "post_date": "05/16/2021 16:23:46",
          "content": "<p>Thank you for the reply :) This is the most important point that I was missing. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1307253": "I am curious how much improvement can be achieved with proper snap-to-grid. My submission LB scores before / after snap-to-grid (which is my final post-processing step) are: **4.457 / 3.806**. You are welcome to share your results if you want.\n\nI am still investigating why snapping points that are predicted to be in the hallways to the training waypoints results in **better performance** than leaving them not post processed. Maybe training waypoints appear also in testing waypoints? Any thoughts on that? :)",
    "1307257": "Waypoints definitely overlap between the train and test sets.",
    "1307390": "Like @tvdwiele mentioned, train and test waypoints overlap quite a bit - my ballpark estimate is something like 80% - 90% of the test waypoints are also in the training data, so that's why it's advantageous to snap to training points.\n\nSnapping used to get me ~0.5m or so, but as my model and other post-processing has gotten better, snapping has become less and less important (though I still do a \"modified snap\" as a final step)",
    "1307934": "The short answer is: Snap to Grid 4.235 -> 4.050\n\nA longer answer is that I'm doing an iterative looping scheme involving different postprocessing steps (cost minimization, leakage, snap-to-grid, but not always in that order), and the score moves up and down with each step as I go around the loops. Once the full loop stops improving the score, then I give up and submit the best score I found. In general, snapping to the training grid gains me about 0.2 and is better (for me) than snapping to the full generated grid.",
    "1310366": "> I am still investigating why snapping points that are predicted to be in the hallways to the training waypoints results in better performance than leaving them not post processed.\n\nThis competition is not about predicting exact x,y coordinates- it is about predicting which pre-defined location was traveled to. As already mentioned, in 80%-90% of the test set we already know these pre-defined locations.\n\n> Maybe training waypoints appear also in testing waypoints?\n\nExactly, This is part of why snapping works so well. I'm surprised at how strong your score is without knowing this already. Great work on your current LB score!\n\nWish I had more free time to have worked on this competition. Lots of fun ideas I would've liked to try.",
    "1310400": "Thank you for the reply :) This is the most important point that I was missing."
  },
  "source": "meta"
}