{
  "id": 240141,
  "title": "Part of the 2nd place solution (Jack's model)",
  "url": "/competitions/indoor-location-navigation/discussion/240141",
  "author_name": "",
  "post_date": "2021-05-18T16:37:51.123075600Z",
  "votes": 36,
  "comment_count": 11,
  "views": 0,
  "content": "<p>My main contribution to the team was to provide the absolute position prediction from signal strength of wifi as an input to the team's ensemble. There are other things I've done in our team, but I'll write about this here. My prediction accounts for about 1/3 of the weight of the ensemble.</p>\n<hr>\n<h1>Data Preparation</h1>\n<ul>\n<li>The spacing of the way points is too sparse to be combined with the wifi information, so the location between way points was interpolated by host's function \"compute_step_positions\" and generate the target value at intervals of 1 second.</li>\n<li>As input features to the model, I generated data with each column having RSSI value for each BSSID, as well as the others. However, I completely ignored the timestamp of the wifi data and instead adopted last_seen_timestamp as the true timestamp.</li>\n<li>The last_seen_timestamp was rounded to the nearest second, and the BSSIDs observed during that time were stored in the same row, and combined with the target.</li>\n</ul>\n<p>I've added some other processing, but roughly, this data is the input for training my model.</p>\n<hr>\n<h1>Model</h1>\n<ul>\n<li>I built one model per floor, because I thought that the information on the other sites, and even on the other floors of the same building, didn't seem to help much in estimating the location.</li>\n<li>What is characteristic, I think, is that I approached the problem of location estimation not as a regression task, but as a multiclass classification task. To be more precise, I discretized the coordinates into square regions of 2m on a side, and then trained a model to predict which region has the highest probability of being present. (For a floor with a width and height of 200 meters, this means a multiclass classification of 10,000 classes.)</li>\n<li>My model is NN (non-RNN), and I introduced my own innovations to make it learn well as a multi-class classification problem.</li>\n<li>By the trained model, the probability of existence in each region for every second was output and the region with the high probability (the average of the x, y values of the top 30 regions) was adopted as the location prediction.</li>\n<li>Depending on the hyperparameters, it took about 4-6h using the GPU in Kaggle Notebook to do a 5-fold CV of all the sites/floors appearing in the test data. (about 1 min per model)</li>\n</ul>\n<hr>\n<h1>Post-processing</h1>\n<ul>\n<li>Since my model is not an RNN, the accuracy is low without post-processing, and it can only make good predictions when <a href=\"https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization\" target=\"_blank\">cost minimization</a> post-processing is applied. Rather, I decided that there is not much benefit to learn the dependencies of time series by RNN, since it can be introduced by the post-processing.</li>\n<li>Since my model is a multiclass classification model, the predictions for each second have a confidence level, and it was effective to weight the predictions by this confidence level in cost minimization.</li>\n<li>It was also effective to iterate the process, ignoring the points where the position changed significantly as a result of the cost minimization, and redoing the cost minimization.</li>\n<li>This, with some minor angular corrections and thinned out to a way point timestamp, was the input to the team's ensemble.</li>\n</ul>\n<hr>\n<h1>Performance</h1>\n<p>The score when this prediction is submitted without ensemble and further post-processing is as follows:</p>\n<p>public : 3.54326<br>\nprivate: 4.07147</p>\n<p>Before I joined the team, I added various further post-processing steps (snap to corridor, snap to grid, and so on) after that, but they were not applied for the input of team's ensemble.<br>\nIf they are applied, the score is:</p>\n<p>public : 2.73972<br>\nprivate: 3.31607</p>\n<p>If I hadn't teamed up and gotten no improvement in my post-processing, the score would have been about this.</p>\n<hr>\n<p>It was a great honor to play in this competition with a very good team.<br>\nThank you to all the team members, the hosts, and the competitors!</p>",
  "messages": [
    {
      "id": "1313656",
      "postDate": "05/18/2021 16:37:51",
      "content": "<p>My main contribution to the team was to provide the absolute position prediction from signal strength of wifi as an input to the team's ensemble. There are other things I've done in our team, but I'll write about this here. My prediction accounts for about 1/3 of the weight of the ensemble.</p>\n<hr>\n<h1>Data Preparation</h1>\n<ul>\n<li>The spacing of the way points is too sparse to be combined with the wifi information, so the location between way points was interpolated by host's function \"compute_step_positions\" and generate the target value at intervals of 1 second.</li>\n<li>As input features to the model, I generated data with each column having RSSI value for each BSSID, as well as the others. However, I completely ignored the timestamp of the wifi data and instead adopted last_seen_timestamp as the true timestamp.</li>\n<li>The last_seen_timestamp was rounded to the nearest second, and the BSSIDs observed during that time were stored in the same row, and combined with the target.</li>\n</ul>\n<p>I've added some other processing, but roughly, this data is the input for training my model.</p>\n<hr>\n<h1>Model</h1>\n<ul>\n<li>I built one model per floor, because I thought that the information on the other sites, and even on the other floors of the same building, didn't seem to help much in estimating the location.</li>\n<li>What is characteristic, I think, is that I approached the problem of location estimation not as a regression task, but as a multiclass classification task. To be more precise, I discretized the coordinates into square regions of 2m on a side, and then trained a model to predict which region has the highest probability of being present. (For a floor with a width and height of 200 meters, this means a multiclass classification of 10,000 classes.)</li>\n<li>My model is NN (non-RNN), and I introduced my own innovations to make it learn well as a multi-class classification problem.</li>\n<li>By the trained model, the probability of existence in each region for every second was output and the region with the high probability (the average of the x, y values of the top 30 regions) was adopted as the location prediction.</li>\n<li>Depending on the hyperparameters, it took about 4-6h using the GPU in Kaggle Notebook to do a 5-fold CV of all the sites/floors appearing in the test data. (about 1 min per model)</li>\n</ul>\n<hr>\n<h1>Post-processing</h1>\n<ul>\n<li>Since my model is not an RNN, the accuracy is low without post-processing, and it can only make good predictions when <a href=\"https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization\" target=\"_blank\">cost minimization</a> post-processing is applied. Rather, I decided that there is not much benefit to learn the dependencies of time series by RNN, since it can be introduced by the post-processing.</li>\n<li>Since my model is a multiclass classification model, the predictions for each second have a confidence level, and it was effective to weight the predictions by this confidence level in cost minimization.</li>\n<li>It was also effective to iterate the process, ignoring the points where the position changed significantly as a result of the cost minimization, and redoing the cost minimization.</li>\n<li>This, with some minor angular corrections and thinned out to a way point timestamp, was the input to the team's ensemble.</li>\n</ul>\n<hr>\n<h1>Performance</h1>\n<p>The score when this prediction is submitted without ensemble and further post-processing is as follows:</p>\n<p>public : 3.54326<br>\nprivate: 4.07147</p>\n<p>Before I joined the team, I added various further post-processing steps (snap to corridor, snap to grid, and so on) after that, but they were not applied for the input of team's ensemble.<br>\nIf they are applied, the score is:</p>\n<p>public : 2.73972<br>\nprivate: 3.31607</p>\n<p>If I hadn't teamed up and gotten no improvement in my post-processing, the score would have been about this.</p>\n<hr>\n<p>It was a great honor to play in this competition with a very good team.<br>\nThank you to all the team members, the hosts, and the competitors!</p>",
      "rawMarkdown": "My main contribution to the team was to provide the absolute position prediction from signal strength of wifi as an input to the team's ensemble. There are other things I've done in our team, but I'll write about this here. My prediction accounts for about 1/3 of the weight of the ensemble.\n\n-----\n\n# Data Preparation\n- The spacing of the way points is too sparse to be combined with the wifi information, so the location between way points was interpolated by host's function \"compute_step_positions\" and generate the target value at intervals of 1 second.\n- As input features to the model, I generated data with each column having RSSI value for each BSSID, as well as the others. However, I completely ignored the timestamp of the wifi data and instead adopted last_seen_timestamp as the true timestamp.\n- The last_seen_timestamp was rounded to the nearest second, and the BSSIDs observed during that time were stored in the same row, and combined with the target.\n\nI've added some other processing, but roughly, this data is the input for training my model.\n\n-----\n\n# Model\n- I built one model per floor, because I thought that the information on the other sites, and even on the other floors of the same building, didn't seem to help much in estimating the location.\n- What is characteristic, I think, is that I approached the problem of location estimation not as a regression task, but as a multiclass classification task. To be more precise, I discretized the coordinates into square regions of 2m on a side, and then trained a model to predict which region has the highest probability of being present. (For a floor with a width and height of 200 meters, this means a multiclass classification of 10,000 classes.)\n- My model is NN (non-RNN), and I introduced my own innovations to make it learn well as a multi-class classification problem.\n- By the trained model, the probability of existence in each region for every second was output and the region with the high probability (the average of the x, y values of the top 30 regions) was adopted as the location prediction.\n- Depending on the hyperparameters, it took about 4-6h using the GPU in Kaggle Notebook to do a 5-fold CV of all the sites/floors appearing in the test data. (about 1 min per model)\n\n-----\n\n# Post-processing\n- Since my model is not an RNN, the accuracy is low without post-processing, and it can only make good predictions when [cost minimization](https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization) post-processing is applied. Rather, I decided that there is not much benefit to learn the dependencies of time series by RNN, since it can be introduced by the post-processing.\n- Since my model is a multiclass classification model, the predictions for each second have a confidence level, and it was effective to weight the predictions by this confidence level in cost minimization.\n- It was also effective to iterate the process, ignoring the points where the position changed significantly as a result of the cost minimization, and redoing the cost minimization.\n- This, with some minor angular corrections and thinned out to a way point timestamp, was the input to the team's ensemble.\n\n-----\n\n# Performance\nThe score when this prediction is submitted without ensemble and further post-processing is as follows:\n\npublic : 3.54326\nprivate: 4.07147\n\nBefore I joined the team, I added various further post-processing steps (snap to corridor, snap to grid, and so on) after that, but they were not applied for the input of team's ensemble.\nIf they are applied, the score is:\n\npublic : 2.73972\nprivate: 3.31607\n\nIf I hadn't teamed up and gotten no improvement in my post-processing, the score would have been about this.\n\n-----\n\nIt was a great honor to play in this competition with a very good team.\nThank you to all the team members, the hosts, and the competitors!",
      "votes": null
    },
    {
      "id": "1313699",
      "postDate": "05/18/2021 17:03:38",
      "content": "<p>Great idea treating it as a multi-class problem!  Were there any tricks you used that you can share for how to train a 10,000 class multi-class prediction network?</p>",
      "rawMarkdown": "Great idea treating it as a multi-class problem!  Were there any tricks you used that you can share for how to train a 10,000 class multi-class prediction network?",
      "votes": null
    },
    {
      "id": "1314114",
      "postDate": "05/19/2021 00:52:10",
      "content": "<p>Thanks for your comment! First of all, I would like to applaud your excellent results in solo.</p>\n<p>The biggest problem in solving this problem as a multiclass classification problem is that regions with close locations are treated as completely different classes. To solve this problem, I devised a structure for the NN so that predictions for classes in close location are highly correlated. This made it possible to train the model much more efficiently than learning for each of a very large number of classes independently.</p>",
      "rawMarkdown": "Thanks for your comment! First of all, I would like to applaud your excellent results in solo.\n\nThe biggest problem in solving this problem as a multiclass classification problem is that regions with close locations are treated as completely different classes. To solve this problem, I devised a structure for the NN so that predictions for classes in close location are highly correlated. This made it possible to train the model much more efficiently than learning for each of a very large number of classes independently.",
      "votes": null
    },
    {
      "id": "1314115",
      "postDate": "05/19/2021 00:53:16",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/rsakata\" target=\"_blank\">@rsakata</a> ! One of many strong finishes.</p>",
      "rawMarkdown": "Congrats @rsakata ! One of many strong finishes.",
      "votes": null
    },
    {
      "id": "1314124",
      "postDate": "05/19/2021 01:02:04",
      "content": "<p>Thanks!</p>\n<p>Neat - I'll have to explore that a bit; I feel like it would be a good tool to have (to know how to train giant multi class problems like that). thanks!</p>",
      "rawMarkdown": "Thanks!\n\nNeat - I'll have to explore that a bit; I feel like it would be a good tool to have (to know how to train giant multi class problems like that). thanks!",
      "votes": null
    },
    {
      "id": "1314216",
      "postDate": "05/19/2021 03:30:56",
      "content": "<p>Thanks for sharing your great solution!<br>\nCould you elaborate on \" I devised a structure for the NN so that predictions for classes in close location are highly correlated. \"?  If possible, I'd like to learn how you did it.</p>\n<p>thank you!</p>",
      "rawMarkdown": "Thanks for sharing your great solution!\nCould you elaborate on \" I devised a structure for the NN so that predictions for classes in close location are highly correlated. \"?  If possible, I'd like to learn how you did it.\n\n\nthank you!",
      "votes": null
    },
    {
      "id": "1314899",
      "postDate": "05/19/2021 12:43:43",
      "content": "<p>Thanks Jack, I really enjoyed the discussion with you and I couldn't improve my discrete optimization process without you!</p>",
      "rawMarkdown": "Thanks Jack, I really enjoyed the discussion with you and I couldn't improve my discrete optimization process without you!",
      "votes": null
    },
    {
      "id": "1315269",
      "postDate": "05/19/2021 17:10:10",
      "content": "<p>In fact, it's a simple trick.<br>\nThe predicted values for each class corresponding to a 2D position were treated like image data, and Gaussian blurring was applied. In other words, I defined a custom GaussianBlur layer and inserted it just before the output layer. This layer spreads the predictions for one region to the surrounding area, and they become correlated.</p>",
      "rawMarkdown": "In fact, it's a simple trick.\nThe predicted values for each class corresponding to a 2D position were treated like image data, and Gaussian blurring was applied. In other words, I defined a custom GaussianBlur layer and inserted it just before the output layer. This layer spreads the predictions for one region to the surrounding area, and they become correlated.",
      "votes": null
    },
    {
      "id": "1315443",
      "postDate": "05/19/2021 19:39:30",
      "content": "<p>Wow, agreed, very interesting to do this as multi-class. Do you also GaussianBlur the ground truth, almost like label smoothing? Or just the prediction? It almost makes it feel a little like an RMSE/MAE regression loss function, because you get penalized less for being close.</p>\n<p>Also, I did a very similar preprocessing, I think relying entirely on lastseen time and constructing your own 1s wifi \"blocks\" was the right way to go.</p>\n<p>Congrats to you and your team!</p>",
      "rawMarkdown": "Wow, agreed, very interesting to do this as multi-class. Do you also GaussianBlur the ground truth, almost like label smoothing? Or just the prediction? It almost makes it feel a little like an RMSE/MAE regression loss function, because you get penalized less for being close.\n\nAlso, I did a very similar preprocessing, I think relying entirely on lastseen time and constructing your own 1s wifi \"blocks\" was the right way to go.\n\nCongrats to you and your team!",
      "votes": null
    },
    {
      "id": "1315555",
      "postDate": "05/19/2021 23:07:32",
      "content": "<p>Yes, you are right! I applied Gaussian blurring only to prediction, just because I wanted to use ordinary cross-entropy loss function. I think either will have the same effect.</p>",
      "rawMarkdown": "Yes, you are right! I applied Gaussian blurring only to prediction, just because I wanted to use ordinary cross-entropy loss function. I think either will have the same effect.",
      "votes": null
    },
    {
      "id": "1315560",
      "postDate": "05/19/2021 23:14:23",
      "content": "<p>Thanks a lot too! It was very exciting to work with you, and I learned a lot of things.</p>",
      "rawMarkdown": "Thanks a lot too! It was very exciting to work with you, and I learned a lot of things.",
      "votes": null
    },
    {
      "id": "1315587",
      "postDate": "05/20/2021 00:38:56",
      "content": "<p>Thank you Jack for the answer. Simple and very smart approach. Thanks for sharing.</p>",
      "rawMarkdown": "Thank you Jack for the answer. Simple and very smart approach. Thanks for sharing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1313699,
      "author_name": "chris62",
      "author_url": "",
      "post_date": "05/18/2021 17:03:38",
      "content": "<p>Great idea treating it as a multi-class problem!  Were there any tricks you used that you can share for how to train a 10,000 class multi-class prediction network?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1314114,
          "author_name": "rsakata",
          "author_url": "",
          "post_date": "05/19/2021 00:52:10",
          "content": "<p>Thanks for your comment! First of all, I would like to applaud your excellent results in solo.</p>\n<p>The biggest problem in solving this problem as a multiclass classification problem is that regions with close locations are treated as completely different classes. To solve this problem, I devised a structure for the NN so that predictions for classes in close location are highly correlated. This made it possible to train the model much more efficiently than learning for each of a very large number of classes independently.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1314124,
          "author_name": "chris62",
          "author_url": "",
          "post_date": "05/19/2021 01:02:04",
          "content": "<p>Thanks!</p>\n<p>Neat - I'll have to explore that a bit; I feel like it would be a good tool to have (to know how to train giant multi class problems like that). thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1314216,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "05/19/2021 03:30:56",
          "content": "<p>Thanks for sharing your great solution!<br>\nCould you elaborate on \" I devised a structure for the NN so that predictions for classes in close location are highly correlated. \"?  If possible, I'd like to learn how you did it.</p>\n<p>thank you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1315269,
          "author_name": "rsakata",
          "author_url": "",
          "post_date": "05/19/2021 17:10:10",
          "content": "<p>In fact, it's a simple trick.<br>\nThe predicted values for each class corresponding to a 2D position were treated like image data, and Gaussian blurring was applied. In other words, I defined a custom GaussianBlur layer and inserted it just before the output layer. This layer spreads the predictions for one region to the surrounding area, and they become correlated.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1315443,
          "author_name": "paulfornia",
          "author_url": "",
          "post_date": "05/19/2021 19:39:30",
          "content": "<p>Wow, agreed, very interesting to do this as multi-class. Do you also GaussianBlur the ground truth, almost like label smoothing? Or just the prediction? It almost makes it feel a little like an RMSE/MAE regression loss function, because you get penalized less for being close.</p>\n<p>Also, I did a very similar preprocessing, I think relying entirely on lastseen time and constructing your own 1s wifi \"blocks\" was the right way to go.</p>\n<p>Congrats to you and your team!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1315555,
          "author_name": "rsakata",
          "author_url": "",
          "post_date": "05/19/2021 23:07:32",
          "content": "<p>Yes, you are right! I applied Gaussian blurring only to prediction, just because I wanted to use ordinary cross-entropy loss function. I think either will have the same effect.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1315587,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "05/20/2021 00:38:56",
          "content": "<p>Thank you Jack for the answer. Simple and very smart approach. Thanks for sharing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1314115,
      "author_name": "robikscube",
      "author_url": "",
      "post_date": "05/19/2021 00:53:16",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/rsakata\" target=\"_blank\">@rsakata</a> ! One of many strong finishes.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1314899,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "05/19/2021 12:43:43",
      "content": "<p>Thanks Jack, I really enjoyed the discussion with you and I couldn't improve my discrete optimization process without you!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1315560,
          "author_name": "rsakata",
          "author_url": "",
          "post_date": "05/19/2021 23:14:23",
          "content": "<p>Thanks a lot too! It was very exciting to work with you, and I learned a lot of things.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1313656": "My main contribution to the team was to provide the absolute position prediction from signal strength of wifi as an input to the team's ensemble. There are other things I've done in our team, but I'll write about this here. My prediction accounts for about 1/3 of the weight of the ensemble.\n\n-----\n\n# Data Preparation\n- The spacing of the way points is too sparse to be combined with the wifi information, so the location between way points was interpolated by host's function \"compute_step_positions\" and generate the target value at intervals of 1 second.\n- As input features to the model, I generated data with each column having RSSI value for each BSSID, as well as the others. However, I completely ignored the timestamp of the wifi data and instead adopted last_seen_timestamp as the true timestamp.\n- The last_seen_timestamp was rounded to the nearest second, and the BSSIDs observed during that time were stored in the same row, and combined with the target.\n\nI've added some other processing, but roughly, this data is the input for training my model.\n\n-----\n\n# Model\n- I built one model per floor, because I thought that the information on the other sites, and even on the other floors of the same building, didn't seem to help much in estimating the location.\n- What is characteristic, I think, is that I approached the problem of location estimation not as a regression task, but as a multiclass classification task. To be more precise, I discretized the coordinates into square regions of 2m on a side, and then trained a model to predict which region has the highest probability of being present. (For a floor with a width and height of 200 meters, this means a multiclass classification of 10,000 classes.)\n- My model is NN (non-RNN), and I introduced my own innovations to make it learn well as a multi-class classification problem.\n- By the trained model, the probability of existence in each region for every second was output and the region with the high probability (the average of the x, y values of the top 30 regions) was adopted as the location prediction.\n- Depending on the hyperparameters, it took about 4-6h using the GPU in Kaggle Notebook to do a 5-fold CV of all the sites/floors appearing in the test data. (about 1 min per model)\n\n-----\n\n# Post-processing\n- Since my model is not an RNN, the accuracy is low without post-processing, and it can only make good predictions when [cost minimization](https://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization) post-processing is applied. Rather, I decided that there is not much benefit to learn the dependencies of time series by RNN, since it can be introduced by the post-processing.\n- Since my model is a multiclass classification model, the predictions for each second have a confidence level, and it was effective to weight the predictions by this confidence level in cost minimization.\n- It was also effective to iterate the process, ignoring the points where the position changed significantly as a result of the cost minimization, and redoing the cost minimization.\n- This, with some minor angular corrections and thinned out to a way point timestamp, was the input to the team's ensemble.\n\n-----\n\n# Performance\nThe score when this prediction is submitted without ensemble and further post-processing is as follows:\n\npublic : 3.54326\nprivate: 4.07147\n\nBefore I joined the team, I added various further post-processing steps (snap to corridor, snap to grid, and so on) after that, but they were not applied for the input of team's ensemble.\nIf they are applied, the score is:\n\npublic : 2.73972\nprivate: 3.31607\n\nIf I hadn't teamed up and gotten no improvement in my post-processing, the score would have been about this.\n\n-----\n\nIt was a great honor to play in this competition with a very good team.\nThank you to all the team members, the hosts, and the competitors!",
    "1313699": "Great idea treating it as a multi-class problem!  Were there any tricks you used that you can share for how to train a 10,000 class multi-class prediction network?",
    "1314114": "Thanks for your comment! First of all, I would like to applaud your excellent results in solo.\n\nThe biggest problem in solving this problem as a multiclass classification problem is that regions with close locations are treated as completely different classes. To solve this problem, I devised a structure for the NN so that predictions for classes in close location are highly correlated. This made it possible to train the model much more efficiently than learning for each of a very large number of classes independently.",
    "1314115": "Congrats @rsakata ! One of many strong finishes.",
    "1314124": "Thanks!\n\nNeat - I'll have to explore that a bit; I feel like it would be a good tool to have (to know how to train giant multi class problems like that). thanks!",
    "1314216": "Thanks for sharing your great solution!\nCould you elaborate on \" I devised a structure for the NN so that predictions for classes in close location are highly correlated. \"?  If possible, I'd like to learn how you did it.\n\n\nthank you!",
    "1314899": "Thanks Jack, I really enjoyed the discussion with you and I couldn't improve my discrete optimization process without you!",
    "1315269": "In fact, it's a simple trick.\nThe predicted values for each class corresponding to a 2D position were treated like image data, and Gaussian blurring was applied. In other words, I defined a custom GaussianBlur layer and inserted it just before the output layer. This layer spreads the predictions for one region to the surrounding area, and they become correlated.",
    "1315443": "Wow, agreed, very interesting to do this as multi-class. Do you also GaussianBlur the ground truth, almost like label smoothing? Or just the prediction? It almost makes it feel a little like an RMSE/MAE regression loss function, because you get penalized less for being close.\n\nAlso, I did a very similar preprocessing, I think relying entirely on lastseen time and constructing your own 1s wifi \"blocks\" was the right way to go.\n\nCongrats to you and your team!",
    "1315555": "Yes, you are right! I applied Gaussian blurring only to prediction, just because I wanted to use ordinary cross-entropy loss function. I think either will have the same effect.",
    "1315560": "Thanks a lot too! It was very exciting to work with you, and I learned a lot of things.",
    "1315587": "Thank you Jack for the answer. Simple and very smart approach. Thanks for sharing."
  },
  "source": "meta"
}