{
  "id": 240058,
  "title": "How did you validate your post processing?",
  "url": "/competitions/indoor-location-navigation/discussion/240058",
  "author_name": "",
  "post_date": "2021-05-18T12:29:01.993110800Z",
  "votes": 12,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hi all,<br>\nNow we all know the importance of post processing. But I'm wondering how do you validate your post processing.</p>\n<p>More concretely say you have post processing A and B, how can you tell which method would have better score and how much would it be better <strong>without</strong> submitting to LB?</p>",
  "messages": [
    {
      "id": "1313155",
      "postDate": "05/18/2021 12:29:01",
      "content": "<p>Hi all,<br>\nNow we all know the importance of post processing. But I'm wondering how do you validate your post processing.</p>\n<p>More concretely say you have post processing A and B, how can you tell which method would have better score and how much would it be better <strong>without</strong> submitting to LB?</p>",
      "rawMarkdown": "Hi all,\nNow we all know the importance of post processing. But I'm wondering how do you validate your post processing.\n\nMore concretely say you have post processing A and B, how can you tell which method would have better score and how much would it be better **without** submitting to LB?",
      "votes": null
    },
    {
      "id": "1313195",
      "postDate": "05/18/2021 12:48:06",
      "content": "<p>I used three primary methods:</p>\n<ol>\n<li><p>I knew my wifi predictions had about 4-5m error on average, so if the post processing took a path too far from that (say, 8m away), then I'd either average back with wifi, or just reset back to the wifi x,y and start over</p></li>\n<li><p>If the post processing distorted the predicted path shape too much (delta x,y between waypoints), then my script would automatically average back to be closer to the predicted shape again</p></li>\n<li><p>I would do 3-4 runs of post processing, and then try to do a \"smart ensemble\", example: if 3 of 4 runs had path centroids near each other and 1 run was far out of alignment, the ensemble would average only the 3 and throw out the (presumably) bad one. Or sometimes, I would sort the runs for each path by which one was closest to the predicted delta x,y, and pick that one, throwing out the others.</p></li>\n</ol>\n<p>Despite all of that, I wish I had a better way to do cross validation on the post processing - I could never really be sure that those 3 methods were working without checking the public LB, and then I was worried about overfitting :)</p>\n<p>I'll try better CV on post-processing next time!</p>",
      "rawMarkdown": "I used three primary methods:\n\n1. I knew my wifi predictions had about 4-5m error on average, so if the post processing took a path too far from that (say, 8m away), then I'd either average back with wifi, or just reset back to the wifi x,y and start over\n\n2. If the post processing distorted the predicted path shape too much (delta x,y between waypoints), then my script would automatically average back to be closer to the predicted shape again\n\n3. I would do 3-4 runs of post processing, and then try to do a \"smart ensemble\", example: if 3 of 4 runs had path centroids near each other and 1 run was far out of alignment, the ensemble would average only the 3 and throw out the (presumably) bad one. Or sometimes, I would sort the runs for each path by which one was closest to the predicted delta x,y, and pick that one, throwing out the others.\n\nDespite all of that, I wish I had a better way to do cross validation on the post processing - I could never really be sure that those 3 methods were working without checking the public LB, and then I was worried about overfitting :)\n\nI'll try better CV on post-processing next time!",
      "votes": null
    },
    {
      "id": "1313305",
      "postDate": "05/18/2021 13:58:10",
      "content": "<p>The txt size of the test set&gt;=2M (look in linux), the time span&gt;=60s, and the number of path points&gt;=5. In the first half of the month,I also selected 626 paths as fake test using above restrictions . No matter whether post-processing is performed or not, my fake test scores are consistent with the lb scores. <br>\n<a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/240077\" target=\"_blank\">link</a></p>",
      "rawMarkdown": "The txt size of the test set>=2M (look in linux), the time span>=60s, and the number of path points>=5. In the first half of the month,I also selected 626 paths as fake test using above restrictions . No matter whether post-processing is performed or not, my fake test scores are consistent with the lb scores. \n[link](https://www.kaggle.com/c/indoor-location-navigation/discussion/240077)",
      "votes": null
    },
    {
      "id": "1313323",
      "postDate": "05/18/2021 14:08:03",
      "content": "<p>I think I did this the \"correct\" way, but it was a bit tedious and not sure how much value it added, essentially pretending that my holdout set was the submission, and the 80% in-fold set of paths was the training data, and stuck to this quite strictly at each step. Here was my process:</p>\n<p>First, I used the provided github step function to get \"ground truth\" x and y values for every step, and then linear interpolation to get ground truth locations at any time stamp I wanted. Different models used different timestamps, so this was important.  (In reading others' solutions, my biggest mistake was putting too much trust in the provided function, I should have modeled these deltas/shapes myself!)</p>\n<p>Next, I did my base modeling, LSTM at granular time steps, and MLP at regular interval (not wifi block times, but conceptually similar). I then repeated for 80% of the paths, with 20% holdout. (I was planning on doing 5-fold, but it would have been very computationally intensive)</p>\n<p>Then, at every step in post processing, I passed in two separate datasets, and spit out two datasets, one for sub, and one for val. Importantly, my val file also carried through those interpolated \"ground truth\" values from the first step. So I was able to easily calculate my mean distance and floor accuracy metrics at any point, even at stages where the granularity was different from the final waypoints. At the point where I interpolated my predictions in time to match the submission data, I similarly interpolated my val set to match actual waypoints, and switched over to using those known waypoints to calculate metrics.</p>\n<p>One tricky part was that several of these post-processing steps also required two separate inputs for sub vs val. For example, for \"snap to grid\", I had two sets of \"training\" waypoints, where the val version was generated from only the out of fold paths. Similarly, I calculated the leaky start/end points separately for the val set, only using the 80%.</p>\n<p>While tedious, this worked pretty well for me, and always seemed pretty correlated to the public LB.</p>",
      "rawMarkdown": "I think I did this the \"correct\" way, but it was a bit tedious and not sure how much value it added, essentially pretending that my holdout set was the submission, and the 80% in-fold set of paths was the training data, and stuck to this quite strictly at each step. Here was my process:\n\nFirst, I used the provided github step function to get \"ground truth\" x and y values for every step, and then linear interpolation to get ground truth locations at any time stamp I wanted. Different models used different timestamps, so this was important.  (In reading others' solutions, my biggest mistake was putting too much trust in the provided function, I should have modeled these deltas/shapes myself!)\n\nNext, I did my base modeling, LSTM at granular time steps, and MLP at regular interval (not wifi block times, but conceptually similar). I then repeated for 80% of the paths, with 20% holdout. (I was planning on doing 5-fold, but it would have been very computationally intensive)\n\nThen, at every step in post processing, I passed in two separate datasets, and spit out two datasets, one for sub, and one for val. Importantly, my val file also carried through those interpolated \"ground truth\" values from the first step. So I was able to easily calculate my mean distance and floor accuracy metrics at any point, even at stages where the granularity was different from the final waypoints. At the point where I interpolated my predictions in time to match the submission data, I similarly interpolated my val set to match actual waypoints, and switched over to using those known waypoints to calculate metrics.\n\nOne tricky part was that several of these post-processing steps also required two separate inputs for sub vs val. For example, for \"snap to grid\", I had two sets of \"training\" waypoints, where the val version was generated from only the out of fold paths. Similarly, I calculated the leaky start/end points separately for the val set, only using the 80%.\n\nWhile tedious, this worked pretty well for me, and always seemed pretty correlated to the public LB.",
      "votes": null
    },
    {
      "id": "1313337",
      "postDate": "05/18/2021 14:12:41",
      "content": "<p>Oh man, this was smart to look at the characteristics of the sub paths. I think it now makes sense why my (and others') cv scores were systematically higher, because we probably have a bunch of short paths that are much harder predict. </p>",
      "rawMarkdown": "Oh man, this was smart to look at the characteristics of the sub paths. I think it now makes sense why my (and others') cv scores were systematically higher, because we probably have a bunch of short paths that are much harder predict.",
      "votes": null
    },
    {
      "id": "1313348",
      "postDate": "05/18/2021 14:15:10",
      "content": "<p>make sense</p>",
      "rawMarkdown": "make sense",
      "votes": null
    },
    {
      "id": "1314062",
      "postDate": "05/18/2021 23:27:34",
      "content": "<p>Thanks for the detailed steps. I'm very impressed by this part \"\"ground truth\" x and y values for every step, and then linear interpolation to get ground truth locations at any time stamp\". I should have done that too.</p>\n<p>I also agreed on the tricky part some post processing requires other inputs like waypoints which complicates the validation step.</p>\n<p>Very well done. Thank you!</p>",
      "rawMarkdown": "Thanks for the detailed steps. I'm very impressed by this part \"\"ground truth\" x and y values for every step, and then linear interpolation to get ground truth locations at any time stamp\". I should have done that too.\n\nI also agreed on the tricky part some post processing requires other inputs like waypoints which complicates the validation step.\n\nVery well done. Thank you!",
      "votes": null
    },
    {
      "id": "1314374",
      "postDate": "05/19/2021 06:18:00",
      "content": "<p>Thanks for sharing.  I should have done this in the early stage of this competition.</p>",
      "rawMarkdown": "Thanks for sharing.  I should have done this in the early stage of this competition.",
      "votes": null
    },
    {
      "id": "1314379",
      "postDate": "05/19/2021 06:20:28",
      "content": "<p>Thank you chris for sharing. You have very smart trick so that post processing wouldn't make it worse!</p>",
      "rawMarkdown": "Thank you chris for sharing. You have very smart trick so that post processing wouldn't make it worse!",
      "votes": null
    },
    {
      "id": "1314951",
      "postDate": "05/19/2021 13:18:29",
      "content": "<p>I just applied postprocessing to my validation set and checked the score, and it is correlated with LB. </p>",
      "rawMarkdown": "I just applied postprocessing to my validation set and checked the score, and it is correlated with LB.",
      "votes": null
    },
    {
      "id": "1315023",
      "postDate": "05/19/2021 13:56:37",
      "content": "<p>thank you ， I think that combining strategy of both you and me will be the host's real train&amp;test split method.</p>",
      "rawMarkdown": "thank you ， I think that combining strategy of both you and me will be the host's real train&test split method.",
      "votes": null
    },
    {
      "id": "1315451",
      "postDate": "05/19/2021 19:45:17",
      "content": "<p>Thanks, and congrats to you! My posts started getting way longer when earlier this week I suddenly and mysteriously found myself with way more free time…</p>",
      "rawMarkdown": "Thanks, and congrats to you! My posts started getting way longer when earlier this week I suddenly and mysteriously found myself with way more free time...",
      "votes": null
    },
    {
      "id": "1315594",
      "postDate": "05/20/2021 01:12:35",
      "content": "<p>This is very interesting question and I want to hear answers of other kagglers.<br>\nI think some post-processing techniques are based on explicit assumptions/hypotheses about test set like <a href=\"https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage\" target=\"_blank\">there would be a leakage</a> or <a href=\"https://www.kaggle.com/robikscube/indoor-navigation-snap-to-grid-post-processing\" target=\"_blank\">train and test waypoints would be overlapped</a>. We can infer these hypotheses throughout EDA of train set or train/val split validation, but it would be difficult to validate it without submitting to LB. These techniques would not work if the host split train/test set in different ways. Of course, <a href=\"https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization\" target=\"_blank\">the elegant solution of your teammate</a> is not the case.</p>",
      "rawMarkdown": "This is very interesting question and I want to hear answers of other kagglers.\nI think some post-processing techniques are based on explicit assumptions/hypotheses about test set like [there would be a leakage](https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage) or [train and test waypoints would be overlapped](https://www.kaggle.com/robikscube/indoor-navigation-snap-to-grid-post-processing). We can infer these hypotheses throughout EDA of train set or train/val split validation, but it would be difficult to validate it without submitting to LB. These techniques would not work if the host split train/test set in different ways. Of course, [the elegant solution of your teammate](https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization) is not the case.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1313195,
      "author_name": "chris62",
      "author_url": "",
      "post_date": "05/18/2021 12:48:06",
      "content": "<p>I used three primary methods:</p>\n<ol>\n<li><p>I knew my wifi predictions had about 4-5m error on average, so if the post processing took a path too far from that (say, 8m away), then I'd either average back with wifi, or just reset back to the wifi x,y and start over</p></li>\n<li><p>If the post processing distorted the predicted path shape too much (delta x,y between waypoints), then my script would automatically average back to be closer to the predicted shape again</p></li>\n<li><p>I would do 3-4 runs of post processing, and then try to do a \"smart ensemble\", example: if 3 of 4 runs had path centroids near each other and 1 run was far out of alignment, the ensemble would average only the 3 and throw out the (presumably) bad one. Or sometimes, I would sort the runs for each path by which one was closest to the predicted delta x,y, and pick that one, throwing out the others.</p></li>\n</ol>\n<p>Despite all of that, I wish I had a better way to do cross validation on the post processing - I could never really be sure that those 3 methods were working without checking the public LB, and then I was worried about overfitting :)</p>\n<p>I'll try better CV on post-processing next time!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1314379,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "05/19/2021 06:20:28",
          "content": "<p>Thank you chris for sharing. You have very smart trick so that post processing wouldn't make it worse!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1313305,
      "author_name": "max2020",
      "author_url": "",
      "post_date": "05/18/2021 13:58:10",
      "content": "<p>The txt size of the test set&gt;=2M (look in linux), the time span&gt;=60s, and the number of path points&gt;=5. In the first half of the month,I also selected 626 paths as fake test using above restrictions . No matter whether post-processing is performed or not, my fake test scores are consistent with the lb scores. <br>\n<a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/240077\" target=\"_blank\">link</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1313337,
          "author_name": "paulfornia",
          "author_url": "",
          "post_date": "05/18/2021 14:12:41",
          "content": "<p>Oh man, this was smart to look at the characteristics of the sub paths. I think it now makes sense why my (and others') cv scores were systematically higher, because we probably have a bunch of short paths that are much harder predict. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1313348,
          "author_name": "max2020",
          "author_url": "",
          "post_date": "05/18/2021 14:15:10",
          "content": "<p>make sense</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1314374,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "05/19/2021 06:18:00",
          "content": "<p>Thanks for sharing.  I should have done this in the early stage of this competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1313323,
      "author_name": "paulfornia",
      "author_url": "",
      "post_date": "05/18/2021 14:08:03",
      "content": "<p>I think I did this the \"correct\" way, but it was a bit tedious and not sure how much value it added, essentially pretending that my holdout set was the submission, and the 80% in-fold set of paths was the training data, and stuck to this quite strictly at each step. Here was my process:</p>\n<p>First, I used the provided github step function to get \"ground truth\" x and y values for every step, and then linear interpolation to get ground truth locations at any time stamp I wanted. Different models used different timestamps, so this was important.  (In reading others' solutions, my biggest mistake was putting too much trust in the provided function, I should have modeled these deltas/shapes myself!)</p>\n<p>Next, I did my base modeling, LSTM at granular time steps, and MLP at regular interval (not wifi block times, but conceptually similar). I then repeated for 80% of the paths, with 20% holdout. (I was planning on doing 5-fold, but it would have been very computationally intensive)</p>\n<p>Then, at every step in post processing, I passed in two separate datasets, and spit out two datasets, one for sub, and one for val. Importantly, my val file also carried through those interpolated \"ground truth\" values from the first step. So I was able to easily calculate my mean distance and floor accuracy metrics at any point, even at stages where the granularity was different from the final waypoints. At the point where I interpolated my predictions in time to match the submission data, I similarly interpolated my val set to match actual waypoints, and switched over to using those known waypoints to calculate metrics.</p>\n<p>One tricky part was that several of these post-processing steps also required two separate inputs for sub vs val. For example, for \"snap to grid\", I had two sets of \"training\" waypoints, where the val version was generated from only the out of fold paths. Similarly, I calculated the leaky start/end points separately for the val set, only using the 80%.</p>\n<p>While tedious, this worked pretty well for me, and always seemed pretty correlated to the public LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1314062,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "05/18/2021 23:27:34",
          "content": "<p>Thanks for the detailed steps. I'm very impressed by this part \"\"ground truth\" x and y values for every step, and then linear interpolation to get ground truth locations at any time stamp\". I should have done that too.</p>\n<p>I also agreed on the tricky part some post processing requires other inputs like waypoints which complicates the validation step.</p>\n<p>Very well done. Thank you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1315451,
          "author_name": "paulfornia",
          "author_url": "",
          "post_date": "05/19/2021 19:45:17",
          "content": "<p>Thanks, and congrats to you! My posts started getting way longer when earlier this week I suddenly and mysteriously found myself with way more free time…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1314951,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "05/19/2021 13:18:29",
      "content": "<p>I just applied postprocessing to my validation set and checked the score, and it is correlated with LB. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1315023,
          "author_name": "max2020",
          "author_url": "",
          "post_date": "05/19/2021 13:56:37",
          "content": "<p>thank you ， I think that combining strategy of both you and me will be the host's real train&amp;test split method.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1315594,
      "author_name": "tomooinubushi",
      "author_url": "",
      "post_date": "05/20/2021 01:12:35",
      "content": "<p>This is very interesting question and I want to hear answers of other kagglers.<br>\nI think some post-processing techniques are based on explicit assumptions/hypotheses about test set like <a href=\"https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage\" target=\"_blank\">there would be a leakage</a> or <a href=\"https://www.kaggle.com/robikscube/indoor-navigation-snap-to-grid-post-processing\" target=\"_blank\">train and test waypoints would be overlapped</a>. We can infer these hypotheses throughout EDA of train set or train/val split validation, but it would be difficult to validate it without submitting to LB. These techniques would not work if the host split train/test set in different ways. Of course, <a href=\"https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization\" target=\"_blank\">the elegant solution of your teammate</a> is not the case.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1313155": "Hi all,\nNow we all know the importance of post processing. But I'm wondering how do you validate your post processing.\n\nMore concretely say you have post processing A and B, how can you tell which method would have better score and how much would it be better **without** submitting to LB?",
    "1313195": "I used three primary methods:\n\n1. I knew my wifi predictions had about 4-5m error on average, so if the post processing took a path too far from that (say, 8m away), then I'd either average back with wifi, or just reset back to the wifi x,y and start over\n\n2. If the post processing distorted the predicted path shape too much (delta x,y between waypoints), then my script would automatically average back to be closer to the predicted shape again\n\n3. I would do 3-4 runs of post processing, and then try to do a \"smart ensemble\", example: if 3 of 4 runs had path centroids near each other and 1 run was far out of alignment, the ensemble would average only the 3 and throw out the (presumably) bad one. Or sometimes, I would sort the runs for each path by which one was closest to the predicted delta x,y, and pick that one, throwing out the others.\n\nDespite all of that, I wish I had a better way to do cross validation on the post processing - I could never really be sure that those 3 methods were working without checking the public LB, and then I was worried about overfitting :)\n\nI'll try better CV on post-processing next time!",
    "1313305": "The txt size of the test set>=2M (look in linux), the time span>=60s, and the number of path points>=5. In the first half of the month,I also selected 626 paths as fake test using above restrictions . No matter whether post-processing is performed or not, my fake test scores are consistent with the lb scores. \n[link](https://www.kaggle.com/c/indoor-location-navigation/discussion/240077)",
    "1313323": "I think I did this the \"correct\" way, but it was a bit tedious and not sure how much value it added, essentially pretending that my holdout set was the submission, and the 80% in-fold set of paths was the training data, and stuck to this quite strictly at each step. Here was my process:\n\nFirst, I used the provided github step function to get \"ground truth\" x and y values for every step, and then linear interpolation to get ground truth locations at any time stamp I wanted. Different models used different timestamps, so this was important.  (In reading others' solutions, my biggest mistake was putting too much trust in the provided function, I should have modeled these deltas/shapes myself!)\n\nNext, I did my base modeling, LSTM at granular time steps, and MLP at regular interval (not wifi block times, but conceptually similar). I then repeated for 80% of the paths, with 20% holdout. (I was planning on doing 5-fold, but it would have been very computationally intensive)\n\nThen, at every step in post processing, I passed in two separate datasets, and spit out two datasets, one for sub, and one for val. Importantly, my val file also carried through those interpolated \"ground truth\" values from the first step. So I was able to easily calculate my mean distance and floor accuracy metrics at any point, even at stages where the granularity was different from the final waypoints. At the point where I interpolated my predictions in time to match the submission data, I similarly interpolated my val set to match actual waypoints, and switched over to using those known waypoints to calculate metrics.\n\nOne tricky part was that several of these post-processing steps also required two separate inputs for sub vs val. For example, for \"snap to grid\", I had two sets of \"training\" waypoints, where the val version was generated from only the out of fold paths. Similarly, I calculated the leaky start/end points separately for the val set, only using the 80%.\n\nWhile tedious, this worked pretty well for me, and always seemed pretty correlated to the public LB.",
    "1313337": "Oh man, this was smart to look at the characteristics of the sub paths. I think it now makes sense why my (and others') cv scores were systematically higher, because we probably have a bunch of short paths that are much harder predict.",
    "1313348": "make sense",
    "1314062": "Thanks for the detailed steps. I'm very impressed by this part \"\"ground truth\" x and y values for every step, and then linear interpolation to get ground truth locations at any time stamp\". I should have done that too.\n\nI also agreed on the tricky part some post processing requires other inputs like waypoints which complicates the validation step.\n\nVery well done. Thank you!",
    "1314374": "Thanks for sharing.  I should have done this in the early stage of this competition.",
    "1314379": "Thank you chris for sharing. You have very smart trick so that post processing wouldn't make it worse!",
    "1314951": "I just applied postprocessing to my validation set and checked the score, and it is correlated with LB.",
    "1315023": "thank you ， I think that combining strategy of both you and me will be the host's real train&test split method.",
    "1315451": "Thanks, and congrats to you! My posts started getting way longer when earlier this week I suddenly and mysteriously found myself with way more free time...",
    "1315594": "This is very interesting question and I want to hear answers of other kagglers.\nI think some post-processing techniques are based on explicit assumptions/hypotheses about test set like [there would be a leakage](https://www.kaggle.com/tomooinubushi/postprocessing-based-on-leakage) or [train and test waypoints would be overlapped](https://www.kaggle.com/robikscube/indoor-navigation-snap-to-grid-post-processing). We can infer these hypotheses throughout EDA of train set or train/val split validation, but it would be difficult to validate it without submitting to LB. These techniques would not work if the host split train/test set in different ways. Of course, [the elegant solution of your teammate](https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization) is not the case."
  },
  "source": "meta"
}