{
  "id": 222580,
  "title": "An example of one dataset for one model",
  "url": "/competitions/indoor-location-navigation/discussion/222580",
  "author_name": "",
  "post_date": "2021-02-28T00:27:19.656257500Z",
  "votes": 47,
  "comment_count": 26,
  "views": 0,
  "content": "<p>I made a WiFi dataset, all the features are the same in all the sites.<br>\nMy current score was observed with this dataset. </p>\n<p><a href=\"https://www.kaggle.com/kokitanisaka/indoorunifiedwifids\" target=\"_blank\">This is the dataset</a>.<br>\nAnd <a href=\"https://www.kaggle.com/kokitanisaka/create-unified-wifi-features-example\" target=\"_blank\">this is the notebook</a> how I made it. <br>\nThe description is in the dataset page.</p>\n<p>The dataset was made from <a href=\"https://www.kaggle.com/hiro5299834\" target=\"_blank\">@hiro5299834</a> 's  <a href=\"https://www.kaggle.com/hiro5299834/indoor-navigation-and-location-wifi-features\" target=\"_blank\">dataset</a>.</p>\n<p>As I wanted to make one model for all the sites, I wanted to put all the data into one frame.<br>\nTo do this, I transformed columns into rows sorting by RSSI strength. </p>\n<p>Hope it helps! </p>\n<p>I think there are many ways to make one dataset through all the sites. <br>\nI appreciate if you can share your ideas. Thank you!</p>\n<p>:Update<br>\nI added .pkl version in the dataset for our convenience. </p>",
  "messages": [
    {
      "id": "1220410",
      "postDate": "02/28/2021 00:27:19",
      "content": "<p>I made a WiFi dataset, all the features are the same in all the sites.<br>\nMy current score was observed with this dataset. </p>\n<p><a href=\"https://www.kaggle.com/kokitanisaka/indoorunifiedwifids\" target=\"_blank\">This is the dataset</a>.<br>\nAnd <a href=\"https://www.kaggle.com/kokitanisaka/create-unified-wifi-features-example\" target=\"_blank\">this is the notebook</a> how I made it. <br>\nThe description is in the dataset page.</p>\n<p>The dataset was made from <a href=\"https://www.kaggle.com/hiro5299834\" target=\"_blank\">@hiro5299834</a> 's  <a href=\"https://www.kaggle.com/hiro5299834/indoor-navigation-and-location-wifi-features\" target=\"_blank\">dataset</a>.</p>\n<p>As I wanted to make one model for all the sites, I wanted to put all the data into one frame.<br>\nTo do this, I transformed columns into rows sorting by RSSI strength. </p>\n<p>Hope it helps! </p>\n<p>I think there are many ways to make one dataset through all the sites. <br>\nI appreciate if you can share your ideas. Thank you!</p>\n<p>:Update<br>\nI added .pkl version in the dataset for our convenience. </p>",
      "rawMarkdown": "I made a WiFi dataset, all the features are the same in all the sites.\nMy current score was observed with this dataset. \n\n[This is the dataset](https://www.kaggle.com/kokitanisaka/indoorunifiedwifids).\nAnd [this is the notebook](https://www.kaggle.com/kokitanisaka/create-unified-wifi-features-example) how I made it. \nThe description is in the dataset page.\n\nThe dataset was made from @hiro5299834 's  [dataset](https://www.kaggle.com/hiro5299834/indoor-navigation-and-location-wifi-features).\n\nAs I wanted to make one model for all the sites, I wanted to put all the data into one frame.\nTo do this, I transformed columns into rows sorting by RSSI strength. \n\nHope it helps! \n\n\nI think there are many ways to make one dataset through all the sites. \nI appreciate if you can share your ideas. Thank you!\n\n:Update\nI added .pkl version in the dataset for our convenience.",
      "votes": null
    },
    {
      "id": "1220420",
      "postDate": "02/28/2021 01:03:27",
      "content": "<p>Thanks for sharing. The link to the notebook is 404 for me. Could you double check if it is working for you?</p>",
      "rawMarkdown": "Thanks for sharing. The link to the notebook is 404 for me. Could you double check if it is working for you?",
      "votes": null
    },
    {
      "id": "1220434",
      "postDate": "02/28/2021 01:51:39",
      "content": "<p>My bad. Thank you for pointing it out!<br>\nI modified the link. Now it would work. </p>",
      "rawMarkdown": "My bad. Thank you for pointing it out!\nI modified the link. Now it would work.",
      "votes": null
    },
    {
      "id": "1220481",
      "postDate": "02/28/2021 03:53:18",
      "content": "<p>It worked for me. Thank you. Great work!</p>",
      "rawMarkdown": "It worked for me. Thank you. Great work!",
      "votes": null
    },
    {
      "id": "1220605",
      "postDate": "02/28/2021 06:40:02",
      "content": "<p>Thank you for sharing your idea! I wanna try it!</p>",
      "rawMarkdown": "Thank you for sharing your idea! I wanna try it!",
      "votes": null
    },
    {
      "id": "1220684",
      "postDate": "02/28/2021 08:49:18",
      "content": "<p>Yes, please try your idea with this! 😄</p>\n<p>I use RNN to make prediction with this dataset. <br>\nI tried LGBM as well, but LGBM wasn't able to learn anything with this dataset during my experiment. Hahaha</p>",
      "rawMarkdown": "Yes, please try your idea with this! 😄\n\nI use RNN to make prediction with this dataset. \nI tried LGBM as well, but LGBM wasn't able to learn anything with this dataset during my experiment. Hahaha",
      "votes": null
    },
    {
      "id": "1220986",
      "postDate": "02/28/2021 15:19:28",
      "content": "<p>Thank you for sharing with the wider community, very helpful!</p>",
      "rawMarkdown": "Thank you for sharing with the wider community, very helpful!",
      "votes": null
    },
    {
      "id": "1221187",
      "postDate": "02/28/2021 18:34:49",
      "content": "<p>Boosting models are not well scalable for multi-tasks problems like this one (both classification and multi_regressions).  You can still have high performance with boosting but training is daunting and models less suited for real world deployement compared to single multi task NN. </p>",
      "rawMarkdown": "Boosting models are not well scalable for multi-tasks problems like this one (both classification and multi_regressions).  You can still have high performance with boosting but training is daunting and models less suited for real world deployement compared to single multi task NN.",
      "votes": null
    },
    {
      "id": "1221970",
      "postDate": "03/01/2021 12:55:00",
      "content": "<p>Thank you for sharing. Do you get your current score using this dataset?</p>",
      "rawMarkdown": "Thank you for sharing. Do you get your current score using this dataset?",
      "votes": null
    },
    {
      "id": "1222471",
      "postDate": "03/01/2021 20:21:19",
      "content": "<p>Yes. For now, it gives me the best result. </p>",
      "rawMarkdown": "Yes. For now, it gives me the best result.",
      "votes": null
    },
    {
      "id": "1222696",
      "postDate": "03/02/2021 04:05:05",
      "content": "<p>Thank you for the source!</p>\n<p>Just in case for reading data frame, we can use:</p>\n<pre><code>df_train = pd.read_csv('../input/indoorunifiedwifids/train_all.csv',  low_memory=False)\n</code></pre>\n<p>I have a little confusion here.<br>\nAt the first, I tried simple models such as lgbm, catboost, but could not work it out though 👀. It also feels like concatenating (over axis 0) all the buildings is against the definition of \"column\"…</p>\n<p>Sorry if I am being too beginner :D , but I'd be appreciated on any advice.</p>",
      "rawMarkdown": "Thank you for the source!\n\nJust in case for reading data frame, we can use:\n```\ndf_train = pd.read_csv('../input/indoorunifiedwifids/train_all.csv',  low_memory=False)\n```\n\nI have a little confusion here.\nAt the first, I tried simple models such as lgbm, catboost, but could not work it out though 👀. It also feels like concatenating (over axis 0) all the buildings is against the definition of \"column\"...\n\nSorry if I am being too beginner :D , but I'd be appreciated on any advice.",
      "votes": null
    },
    {
      "id": "1222721",
      "postDate": "03/02/2021 04:55:11",
      "content": "<p>You're right. LGBM and Catboost won't work with this dataset.<br>\nAs you say, these features don't make sense one feature by one feature. <br>\nInteractions of these columns are important. <br>\nTo utilize this dataset, NeuralNet is the only one option that I can come up with.</p>\n<p>And to make these features sense, Embedding layer or Transformer should be applied. </p>",
      "rawMarkdown": "You're right. LGBM and Catboost won't work with this dataset.\nAs you say, these features don't make sense one feature by one feature. \nInteractions of these columns are important. \nTo utilize this dataset, NeuralNet is the only one option that I can come up with.\n\nAnd to make these features sense, Embedding layer or Transformer should be applied.",
      "votes": null
    },
    {
      "id": "1225805",
      "postDate": "03/04/2021 00:37:02",
      "content": "<p>If you don't mind could you share a bit more with us?<br>\nBecause you mentioned Embedding and Transformer. Are you using RNN model (or variant) and treat the data as time series? Because bssids and rssi are sorted. I think we could use them as time series data. But I wasn't entirely sure.</p>",
      "rawMarkdown": "If you don't mind could you share a bit more with us?\nBecause you mentioned Embedding and Transformer. Are you using RNN model (or variant) and treat the data as time series? Because bssids and rssi are sorted. I think we could use them as time series data. But I wasn't entirely sure.",
      "votes": null
    },
    {
      "id": "1225822",
      "postDate": "03/04/2021 01:35:52",
      "content": "<p>I won't say that I treat the dataset as timeseries as the data doesn't have \"time\" related column.<br>\nIn my understanding, I use RNN (more specifically, currently using LSTM) to make these columns sense. </p>\n<p>The meaning of \"a row\" of the dataset is, the device(cell phone) was observed from multiple wifi access points when the guy was there. The nearest wifi access point is bssid_0, the second nearest one is bssid_1 and so on. And the rssi of bssid_0 is rssi_0, rssi of bssid_1 is rssi_1. </p>\n<p>This is the reason that I said \"interaction of these columns are important\". If the model doesn't understand the relationship between these columns, the model can't learn with this dataset. </p>\n<p>I mentioned Embedding because bssids are categorical features. And they are different from site to site, one-hot-encoding kinda thing isn't realistic. Also, Embedding have ability to express the closeness of categorical values. <br>\nAnd as for Transformer, especially Encoder can be applied for this matter, in my understanding. To make the model learn \"the interaction between the columns\", Transformer can handle the issue. <br>\nLike natural language processing,  Transformer can learn the \"closeness\" of words according to context, am I right? I was thinking we can treat bssid features similar to text. </p>\n<p>If I'm saying something dumb, please correct me. Thank you!</p>",
      "rawMarkdown": "I won't say that I treat the dataset as timeseries as the data doesn't have \"time\" related column.\nIn my understanding, I use RNN (more specifically, currently using LSTM) to make these columns sense. \n\nThe meaning of \"a row\" of the dataset is, the device(cell phone) was observed from multiple wifi access points when the guy was there. The nearest wifi access point is bssid_0, the second nearest one is bssid_1 and so on. And the rssi of bssid_0 is rssi_0, rssi of bssid_1 is rssi_1. \n\nThis is the reason that I said \"interaction of these columns are important\". If the model doesn't understand the relationship between these columns, the model can't learn with this dataset. \n\nI mentioned Embedding because bssids are categorical features. And they are different from site to site, one-hot-encoding kinda thing isn't realistic. Also, Embedding have ability to express the closeness of categorical values. \nAnd as for Transformer, especially Encoder can be applied for this matter, in my understanding. To make the model learn \"the interaction between the columns\", Transformer can handle the issue. \nLike natural language processing,  Transformer can learn the \"closeness\" of words according to context, am I right? I was thinking we can treat bssid features similar to text. \n\n\nIf I'm saying something dumb, please correct me. Thank you!",
      "votes": null
    },
    {
      "id": "1225871",
      "postDate": "03/04/2021 02:54:10",
      "content": "<p>What you mentioned makes sense to me and it was my understanding too. Thank you for confirming.<br>\nbtw are you having two separate LSTMs for rssi_n and bssid_n?</p>\n<p>FYI: I mentioned time series just because in RNN/LSTM context they sometimes refer nth input as item at time=t. In you case bssid_t where (t=[0, …, 99]).</p>\n<p>thank you!</p>",
      "rawMarkdown": "What you mentioned makes sense to me and it was my understanding too. Thank you for confirming.\nbtw are you having two separate LSTMs for rssi_n and bssid_n?\n\nFYI: I mentioned time series just because in RNN/LSTM context they sometimes refer nth input as item at time=t. In you case bssid_t where (t=[0, ..., 99]).\n\nthank you!",
      "votes": null
    },
    {
      "id": "1225877",
      "postDate": "03/04/2021 02:59:30",
      "content": "<p>Great, thanks! 😄</p>\n<p>I can't say which one is better but I'm using only one LSTM layer for bssids and rssis.<br>\nIn other words, I haven't tried 2 separated LSTMs yet, or haven't thought about it.</p>",
      "rawMarkdown": "Great, thanks! 😄\n\nI can't say which one is better but I'm using only one LSTM layer for bssids and rssis.\nIn other words, I haven't tried 2 separated LSTMs yet, or haven't thought about it.",
      "votes": null
    },
    {
      "id": "1226049",
      "postDate": "03/04/2021 07:17:09",
      "content": "<p>Thank you for your reply. Now I have better understanding of your work. Thanks again. I'll experiment with it.</p>",
      "rawMarkdown": "Thank you for your reply. Now I have better understanding of your work. Thanks again. I'll experiment with it.",
      "votes": null
    },
    {
      "id": "1227400",
      "postDate": "03/05/2021 13:49:58",
      "content": "<p>Hi Kouki, Thank you for sharing your great dataset. I'm really impressed with this. <br>\nI have one question about your dataset. <br>\nIn my understanding, TYPE_WIFI and TYPE_WAYPOINT usually have different timestamps.<br>\nSo, the way to assign bssid_* and rssi_* for each x, y looks very difficult for me. How did you do that?</p>",
      "rawMarkdown": "Hi Kouki, Thank you for sharing your great dataset. I'm really impressed with this. \nI have one question about your dataset. \nIn my understanding, TYPE_WIFI and TYPE_WAYPOINT usually have different timestamps.\nSo, the way to assign bssid_* and rssi_* for each x, y looks very difficult for me. How did you do that?",
      "votes": null
    },
    {
      "id": "1227421",
      "postDate": "03/05/2021 14:06:52",
      "content": "<p>The original wifi dataset does this. It just matches closest wifi timestamp group with an waypoint.<br>\nSo obviously there's a room to improve timestamp delta between them.</p>",
      "rawMarkdown": "The original wifi dataset does this. It just matches closest wifi timestamp group with an waypoint.\nSo obviously there's a room to improve timestamp delta between them.",
      "votes": null
    },
    {
      "id": "1227424",
      "postDate": "03/05/2021 14:11:05",
      "content": "<p>Thank you🙏🙏🙏 this idea looks reasonable!</p>",
      "rawMarkdown": "Thank you🙏🙏🙏 this idea looks reasonable!",
      "votes": null
    },
    {
      "id": "1227770",
      "postDate": "03/05/2021 20:55:30",
      "content": "<p>I'm glad that you found it useful, <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> !</p>\n<p>As <a href=\"https://www.kaggle.com/higepon\" target=\"_blank\">@higepon</a> says, it was there in the original dataset.<br>\nThe way they made the dataset can be found in <a href=\"https://www.kaggle.com/devinanzelmo\" target=\"_blank\">@devinanzelmo</a> 's <a href=\"https://www.kaggle.com/devinanzelmo/wifi-features\" target=\"_blank\">notebook</a>.<br>\nAnd <a href=\"https://www.kaggle.com/higepon\" target=\"_blank\">@higepon</a> showed <a href=\"https://www.kaggle.com/higepon/generate-wifi-features-5-times-faster\" target=\"_blank\">the faster version</a> of it. </p>",
      "rawMarkdown": "I'm glad that you found it useful, @mamasinkgs !\n\nAs @higepon says, it was there in the original dataset.\nThe way they made the dataset can be found in @devinanzelmo 's [notebook](https://www.kaggle.com/devinanzelmo/wifi-features).\nAnd @higepon showed [the faster version](https://www.kaggle.com/higepon/generate-wifi-features-5-times-faster) of it.",
      "votes": null
    },
    {
      "id": "1236346",
      "postDate": "03/13/2021 03:14:13",
      "content": "<p>As a beginner, really learned a lot from your notebook. Thx a lot.</p>",
      "rawMarkdown": "As a beginner, really learned a lot from your notebook. Thx a lot.",
      "votes": null
    },
    {
      "id": "1236353",
      "postDate": "03/13/2021 03:29:53",
      "content": "<p>Good to hear that, I'm glad about it. 😄<br>\nLet's enjoy the comp! </p>",
      "rawMarkdown": "Good to hear that, I'm glad about it. 😄\nLet's enjoy the comp!",
      "votes": null
    },
    {
      "id": "1266772",
      "postDate": "04/08/2021 04:26:37",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kokitanisaka\" target=\"_blank\">@kokitanisaka</a> I have a question for you. What happen with unseen categorical values at inference? </p>\n<p>For example, lets say that for a test sample the column <code>bssid_0</code> has a bssid string never seen on train (it might appear on <code>bssid_1</code>, …, <code>bssid_N</code>, but not in <code>bssid_0</code>). Would the model be able to create a meaningful encoding in such situation?</p>",
      "rawMarkdown": "Hi @kokitanisaka I have a question for you. What happen with unseen categorical values at inference? \n\nFor example, lets say that for a test sample the column `bssid_0` has a bssid string never seen on train (it might appear on `bssid_1`, ..., `bssid_N`, but not in `bssid_0`). Would the model be able to create a meaningful encoding in such situation?",
      "votes": null
    },
    {
      "id": "1267001",
      "postDate": "04/08/2021 08:29:10",
      "content": "<p>Thank you for the good question, <a href=\"https://www.kaggle.com/mavillan\" target=\"_blank\">@mavillan</a> .<br>\nTo make it right, I put <code>Embedding</code> layer. I expect it learning the closeness of these bssids. <br>\nEven though a bssid which is not observed in bssid_0 but bssid_1, I think it can figure out how to deal with it. </p>\n<p>But of course, it can't know how to deal with a bssid only appears in the test set. In that case, I don't think it can make a right prediction using bssids. </p>",
      "rawMarkdown": "Thank you for the good question, @mavillan .\nTo make it right, I put `Embedding` layer. I expect it learning the closeness of these bssids. \nEven though a bssid which is not observed in bssid_0 but bssid_1, I think it can figure out how to deal with it. \n\nBut of course, it can't know how to deal with a bssid only appears in the test set. In that case, I don't think it can make a right prediction using bssids.",
      "votes": null
    },
    {
      "id": "1267505",
      "postDate": "04/08/2021 15:04:53",
      "content": "<p>Got it, thanks!  I think we need to filter bssids that have not be seen on train step to avoid problems</p>",
      "rawMarkdown": "Got it, thanks!  I think we need to filter bssids that have not be seen on train step to avoid problems",
      "votes": null
    },
    {
      "id": "1291730",
      "postDate": "05/03/2021 09:40:20",
      "content": "<p>Thank you for good information.</p>",
      "rawMarkdown": "Thank you for good information.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1220420,
      "author_name": "higepon",
      "author_url": "",
      "post_date": "02/28/2021 01:03:27",
      "content": "<p>Thanks for sharing. The link to the notebook is 404 for me. Could you double check if it is working for you?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1220434,
          "author_name": "kokitanisaka",
          "author_url": "",
          "post_date": "02/28/2021 01:51:39",
          "content": "<p>My bad. Thank you for pointing it out!<br>\nI modified the link. Now it would work. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1220481,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "02/28/2021 03:53:18",
          "content": "<p>It worked for me. Thank you. Great work!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1220605,
      "author_name": "res1235",
      "author_url": "",
      "post_date": "02/28/2021 06:40:02",
      "content": "<p>Thank you for sharing your idea! I wanna try it!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1220684,
          "author_name": "kokitanisaka",
          "author_url": "",
          "post_date": "02/28/2021 08:49:18",
          "content": "<p>Yes, please try your idea with this! 😄</p>\n<p>I use RNN to make prediction with this dataset. <br>\nI tried LGBM as well, but LGBM wasn't able to learn anything with this dataset during my experiment. Hahaha</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1221187,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "02/28/2021 18:34:49",
          "content": "<p>Boosting models are not well scalable for multi-tasks problems like this one (both classification and multi_regressions).  You can still have high performance with boosting but training is daunting and models less suited for real world deployement compared to single multi task NN. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1220986,
      "author_name": "saurabhbagchi",
      "author_url": "",
      "post_date": "02/28/2021 15:19:28",
      "content": "<p>Thank you for sharing with the wider community, very helpful!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1221970,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "03/01/2021 12:55:00",
      "content": "<p>Thank you for sharing. Do you get your current score using this dataset?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1222471,
          "author_name": "kokitanisaka",
          "author_url": "",
          "post_date": "03/01/2021 20:21:19",
          "content": "<p>Yes. For now, it gives me the best result. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1222696,
      "author_name": "bayartsogtya",
      "author_url": "",
      "post_date": "03/02/2021 04:05:05",
      "content": "<p>Thank you for the source!</p>\n<p>Just in case for reading data frame, we can use:</p>\n<pre><code>df_train = pd.read_csv('../input/indoorunifiedwifids/train_all.csv',  low_memory=False)\n</code></pre>\n<p>I have a little confusion here.<br>\nAt the first, I tried simple models such as lgbm, catboost, but could not work it out though 👀. It also feels like concatenating (over axis 0) all the buildings is against the definition of \"column\"…</p>\n<p>Sorry if I am being too beginner :D , but I'd be appreciated on any advice.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1222721,
          "author_name": "kokitanisaka",
          "author_url": "",
          "post_date": "03/02/2021 04:55:11",
          "content": "<p>You're right. LGBM and Catboost won't work with this dataset.<br>\nAs you say, these features don't make sense one feature by one feature. <br>\nInteractions of these columns are important. <br>\nTo utilize this dataset, NeuralNet is the only one option that I can come up with.</p>\n<p>And to make these features sense, Embedding layer or Transformer should be applied. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1225805,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "03/04/2021 00:37:02",
          "content": "<p>If you don't mind could you share a bit more with us?<br>\nBecause you mentioned Embedding and Transformer. Are you using RNN model (or variant) and treat the data as time series? Because bssids and rssi are sorted. I think we could use them as time series data. But I wasn't entirely sure.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1225822,
          "author_name": "kokitanisaka",
          "author_url": "",
          "post_date": "03/04/2021 01:35:52",
          "content": "<p>I won't say that I treat the dataset as timeseries as the data doesn't have \"time\" related column.<br>\nIn my understanding, I use RNN (more specifically, currently using LSTM) to make these columns sense. </p>\n<p>The meaning of \"a row\" of the dataset is, the device(cell phone) was observed from multiple wifi access points when the guy was there. The nearest wifi access point is bssid_0, the second nearest one is bssid_1 and so on. And the rssi of bssid_0 is rssi_0, rssi of bssid_1 is rssi_1. </p>\n<p>This is the reason that I said \"interaction of these columns are important\". If the model doesn't understand the relationship between these columns, the model can't learn with this dataset. </p>\n<p>I mentioned Embedding because bssids are categorical features. And they are different from site to site, one-hot-encoding kinda thing isn't realistic. Also, Embedding have ability to express the closeness of categorical values. <br>\nAnd as for Transformer, especially Encoder can be applied for this matter, in my understanding. To make the model learn \"the interaction between the columns\", Transformer can handle the issue. <br>\nLike natural language processing,  Transformer can learn the \"closeness\" of words according to context, am I right? I was thinking we can treat bssid features similar to text. </p>\n<p>If I'm saying something dumb, please correct me. Thank you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1225871,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "03/04/2021 02:54:10",
          "content": "<p>What you mentioned makes sense to me and it was my understanding too. Thank you for confirming.<br>\nbtw are you having two separate LSTMs for rssi_n and bssid_n?</p>\n<p>FYI: I mentioned time series just because in RNN/LSTM context they sometimes refer nth input as item at time=t. In you case bssid_t where (t=[0, …, 99]).</p>\n<p>thank you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1225877,
          "author_name": "kokitanisaka",
          "author_url": "",
          "post_date": "03/04/2021 02:59:30",
          "content": "<p>Great, thanks! 😄</p>\n<p>I can't say which one is better but I'm using only one LSTM layer for bssids and rssis.<br>\nIn other words, I haven't tried 2 separated LSTMs yet, or haven't thought about it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1226049,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "03/04/2021 07:17:09",
          "content": "<p>Thank you for your reply. Now I have better understanding of your work. Thanks again. I'll experiment with it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1227400,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "03/05/2021 13:49:58",
      "content": "<p>Hi Kouki, Thank you for sharing your great dataset. I'm really impressed with this. <br>\nI have one question about your dataset. <br>\nIn my understanding, TYPE_WIFI and TYPE_WAYPOINT usually have different timestamps.<br>\nSo, the way to assign bssid_* and rssi_* for each x, y looks very difficult for me. How did you do that?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1227421,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "03/05/2021 14:06:52",
          "content": "<p>The original wifi dataset does this. It just matches closest wifi timestamp group with an waypoint.<br>\nSo obviously there's a room to improve timestamp delta between them.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1227424,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "03/05/2021 14:11:05",
          "content": "<p>Thank you🙏🙏🙏 this idea looks reasonable!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1227770,
          "author_name": "kokitanisaka",
          "author_url": "",
          "post_date": "03/05/2021 20:55:30",
          "content": "<p>I'm glad that you found it useful, <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> !</p>\n<p>As <a href=\"https://www.kaggle.com/higepon\" target=\"_blank\">@higepon</a> says, it was there in the original dataset.<br>\nThe way they made the dataset can be found in <a href=\"https://www.kaggle.com/devinanzelmo\" target=\"_blank\">@devinanzelmo</a> 's <a href=\"https://www.kaggle.com/devinanzelmo/wifi-features\" target=\"_blank\">notebook</a>.<br>\nAnd <a href=\"https://www.kaggle.com/higepon\" target=\"_blank\">@higepon</a> showed <a href=\"https://www.kaggle.com/higepon/generate-wifi-features-5-times-faster\" target=\"_blank\">the faster version</a> of it. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1236346,
      "author_name": "taolearnstolearn",
      "author_url": "",
      "post_date": "03/13/2021 03:14:13",
      "content": "<p>As a beginner, really learned a lot from your notebook. Thx a lot.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1236353,
          "author_name": "kokitanisaka",
          "author_url": "",
          "post_date": "03/13/2021 03:29:53",
          "content": "<p>Good to hear that, I'm glad about it. 😄<br>\nLet's enjoy the comp! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1266772,
      "author_name": "mavillan",
      "author_url": "",
      "post_date": "04/08/2021 04:26:37",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kokitanisaka\" target=\"_blank\">@kokitanisaka</a> I have a question for you. What happen with unseen categorical values at inference? </p>\n<p>For example, lets say that for a test sample the column <code>bssid_0</code> has a bssid string never seen on train (it might appear on <code>bssid_1</code>, …, <code>bssid_N</code>, but not in <code>bssid_0</code>). Would the model be able to create a meaningful encoding in such situation?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1267001,
          "author_name": "kokitanisaka",
          "author_url": "",
          "post_date": "04/08/2021 08:29:10",
          "content": "<p>Thank you for the good question, <a href=\"https://www.kaggle.com/mavillan\" target=\"_blank\">@mavillan</a> .<br>\nTo make it right, I put <code>Embedding</code> layer. I expect it learning the closeness of these bssids. <br>\nEven though a bssid which is not observed in bssid_0 but bssid_1, I think it can figure out how to deal with it. </p>\n<p>But of course, it can't know how to deal with a bssid only appears in the test set. In that case, I don't think it can make a right prediction using bssids. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1267505,
          "author_name": "mavillan",
          "author_url": "",
          "post_date": "04/08/2021 15:04:53",
          "content": "<p>Got it, thanks!  I think we need to filter bssids that have not be seen on train step to avoid problems</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1291730,
      "author_name": "lys620",
      "author_url": "",
      "post_date": "05/03/2021 09:40:20",
      "content": "<p>Thank you for good information.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1220410": "I made a WiFi dataset, all the features are the same in all the sites.\nMy current score was observed with this dataset. \n\n[This is the dataset](https://www.kaggle.com/kokitanisaka/indoorunifiedwifids).\nAnd [this is the notebook](https://www.kaggle.com/kokitanisaka/create-unified-wifi-features-example) how I made it. \nThe description is in the dataset page.\n\nThe dataset was made from @hiro5299834 's  [dataset](https://www.kaggle.com/hiro5299834/indoor-navigation-and-location-wifi-features).\n\nAs I wanted to make one model for all the sites, I wanted to put all the data into one frame.\nTo do this, I transformed columns into rows sorting by RSSI strength. \n\nHope it helps! \n\n\nI think there are many ways to make one dataset through all the sites. \nI appreciate if you can share your ideas. Thank you!\n\n:Update\nI added .pkl version in the dataset for our convenience.",
    "1220420": "Thanks for sharing. The link to the notebook is 404 for me. Could you double check if it is working for you?",
    "1220434": "My bad. Thank you for pointing it out!\nI modified the link. Now it would work.",
    "1220481": "It worked for me. Thank you. Great work!",
    "1220605": "Thank you for sharing your idea! I wanna try it!",
    "1220684": "Yes, please try your idea with this! 😄\n\nI use RNN to make prediction with this dataset. \nI tried LGBM as well, but LGBM wasn't able to learn anything with this dataset during my experiment. Hahaha",
    "1220986": "Thank you for sharing with the wider community, very helpful!",
    "1221187": "Boosting models are not well scalable for multi-tasks problems like this one (both classification and multi_regressions).  You can still have high performance with boosting but training is daunting and models less suited for real world deployement compared to single multi task NN.",
    "1221970": "Thank you for sharing. Do you get your current score using this dataset?",
    "1222471": "Yes. For now, it gives me the best result.",
    "1222696": "Thank you for the source!\n\nJust in case for reading data frame, we can use:\n```\ndf_train = pd.read_csv('../input/indoorunifiedwifids/train_all.csv',  low_memory=False)\n```\n\nI have a little confusion here.\nAt the first, I tried simple models such as lgbm, catboost, but could not work it out though 👀. It also feels like concatenating (over axis 0) all the buildings is against the definition of \"column\"...\n\nSorry if I am being too beginner :D , but I'd be appreciated on any advice.",
    "1222721": "You're right. LGBM and Catboost won't work with this dataset.\nAs you say, these features don't make sense one feature by one feature. \nInteractions of these columns are important. \nTo utilize this dataset, NeuralNet is the only one option that I can come up with.\n\nAnd to make these features sense, Embedding layer or Transformer should be applied.",
    "1225805": "If you don't mind could you share a bit more with us?\nBecause you mentioned Embedding and Transformer. Are you using RNN model (or variant) and treat the data as time series? Because bssids and rssi are sorted. I think we could use them as time series data. But I wasn't entirely sure.",
    "1225822": "I won't say that I treat the dataset as timeseries as the data doesn't have \"time\" related column.\nIn my understanding, I use RNN (more specifically, currently using LSTM) to make these columns sense. \n\nThe meaning of \"a row\" of the dataset is, the device(cell phone) was observed from multiple wifi access points when the guy was there. The nearest wifi access point is bssid_0, the second nearest one is bssid_1 and so on. And the rssi of bssid_0 is rssi_0, rssi of bssid_1 is rssi_1. \n\nThis is the reason that I said \"interaction of these columns are important\". If the model doesn't understand the relationship between these columns, the model can't learn with this dataset. \n\nI mentioned Embedding because bssids are categorical features. And they are different from site to site, one-hot-encoding kinda thing isn't realistic. Also, Embedding have ability to express the closeness of categorical values. \nAnd as for Transformer, especially Encoder can be applied for this matter, in my understanding. To make the model learn \"the interaction between the columns\", Transformer can handle the issue. \nLike natural language processing,  Transformer can learn the \"closeness\" of words according to context, am I right? I was thinking we can treat bssid features similar to text. \n\n\nIf I'm saying something dumb, please correct me. Thank you!",
    "1225871": "What you mentioned makes sense to me and it was my understanding too. Thank you for confirming.\nbtw are you having two separate LSTMs for rssi_n and bssid_n?\n\nFYI: I mentioned time series just because in RNN/LSTM context they sometimes refer nth input as item at time=t. In you case bssid_t where (t=[0, ..., 99]).\n\nthank you!",
    "1225877": "Great, thanks! 😄\n\nI can't say which one is better but I'm using only one LSTM layer for bssids and rssis.\nIn other words, I haven't tried 2 separated LSTMs yet, or haven't thought about it.",
    "1226049": "Thank you for your reply. Now I have better understanding of your work. Thanks again. I'll experiment with it.",
    "1227400": "Hi Kouki, Thank you for sharing your great dataset. I'm really impressed with this. \nI have one question about your dataset. \nIn my understanding, TYPE_WIFI and TYPE_WAYPOINT usually have different timestamps.\nSo, the way to assign bssid_* and rssi_* for each x, y looks very difficult for me. How did you do that?",
    "1227421": "The original wifi dataset does this. It just matches closest wifi timestamp group with an waypoint.\nSo obviously there's a room to improve timestamp delta between them.",
    "1227424": "Thank you🙏🙏🙏 this idea looks reasonable!",
    "1227770": "I'm glad that you found it useful, @mamasinkgs !\n\nAs @higepon says, it was there in the original dataset.\nThe way they made the dataset can be found in @devinanzelmo 's [notebook](https://www.kaggle.com/devinanzelmo/wifi-features).\nAnd @higepon showed [the faster version](https://www.kaggle.com/higepon/generate-wifi-features-5-times-faster) of it.",
    "1236346": "As a beginner, really learned a lot from your notebook. Thx a lot.",
    "1236353": "Good to hear that, I'm glad about it. 😄\nLet's enjoy the comp!",
    "1266772": "Hi @kokitanisaka I have a question for you. What happen with unseen categorical values at inference? \n\nFor example, lets say that for a test sample the column `bssid_0` has a bssid string never seen on train (it might appear on `bssid_1`, ..., `bssid_N`, but not in `bssid_0`). Would the model be able to create a meaningful encoding in such situation?",
    "1267001": "Thank you for the good question, @mavillan .\nTo make it right, I put `Embedding` layer. I expect it learning the closeness of these bssids. \nEven though a bssid which is not observed in bssid_0 but bssid_1, I think it can figure out how to deal with it. \n\nBut of course, it can't know how to deal with a bssid only appears in the test set. In that case, I don't think it can make a right prediction using bssids.",
    "1267505": "Got it, thanks!  I think we need to filter bssids that have not be seen on train step to avoid problems",
    "1291730": "Thank you for good information."
  },
  "source": "meta"
}