{
  "id": 236096,
  "title": "Tricks for using the sensor data",
  "url": "/competitions/indoor-location-navigation/discussion/236096",
  "author_name": "",
  "post_date": "2021-05-02T20:54:53.627752300Z",
  "votes": 52,
  "comment_count": 8,
  "views": 0,
  "content": "<p>If you're having trouble getting past the 4-5m mark and haven't looked at the sensor data yet, here are a few tricks to get you started:</p>\n<p><strong>1. There's too much sensor data</strong></p>\n<p>The longest test path has 36,577 sensor points, each with 12 values - that's a lot of data!</p>\n<p>The key for me was to figure out how to reduce that count. Averaging can work well (and there are other methods as well). For example - here's a path with 3 waypoints with all of the acceleration data, and then that data averaged into 100 points: <a href=\"https://imgur.com/a/sTZsp3i\" target=\"_blank\">https://imgur.com/a/sTZsp3i</a> (kaggle isn't allowing file uploads right now, so I had to use imgur).</p>\n<p>The key to figure out the right level of data reduction is to do EDA (exploratory data analysis), and a lot of graphing; and then a way to feed that into my model.</p>\n<p><strong>2. The \"z\" axis is the most important</strong></p>\n<p>The project github link (<a href=\"https://github.com/location-competition/indoor-location-competition-20\" target=\"_blank\">https://github.com/location-competition/indoor-location-competition-20</a>) has a link to the Android sensor data page (<a href=\"https://developer.android.com/guide/topics/sensors/sensors_overview\" target=\"_blank\">https://developer.android.com/guide/topics/sensors/sensors_overview</a>) which has an image describing the sensor coordinate system. When you combine that with the description of how the data was collected (the phone was held flat in front of the user), it's clear that the \"z\" axis will be the most important for the sensor values (based on how people move and rotate if they are holding a phone flat in front of them).</p>\n<p>In my own tests, I was able to see about 80% of the value by just using the z axis - so if you're stuck with too much data, start there.</p>\n<p><strong>3. RNNs are difficult to train</strong></p>\n<p>I had a lot of trouble getting an RNN or LSTM to train on the sensor data - but there's more than one way to handle a time series! You can start with just a regular feed forward neural net, a CNN (my favorite), a transformer, or even another ML method like xgboost if you engineer features first.</p>\n<p>Don't get stuck thinking you need a recurrent network just because the popular notebooks use one - you can handle this data in a number of different ways to get value from it.</p>\n<p>You may also think that you need to understand Kalman filters (or similar) before you can use sensor data - but not so! You can start using sensor data in neural nets (for example) without understanding it at all :)</p>\n<p><strong>Overall</strong></p>\n<p>I would start by graphing a bunch of the sensor data yourself - you'll start to see patterns in the paths, around the waypoints, etc, and that should guide you with how to use it.</p>\n<p>I hope this helps some!</p>\n<p>I couldn't get past ~4.5m with just wifi, so hopefully adding some sensor data will help you too :)</p>",
  "messages": [
    {
      "id": "1291248",
      "postDate": "05/02/2021 20:54:53",
      "content": "<p>If you're having trouble getting past the 4-5m mark and haven't looked at the sensor data yet, here are a few tricks to get you started:</p>\n<p><strong>1. There's too much sensor data</strong></p>\n<p>The longest test path has 36,577 sensor points, each with 12 values - that's a lot of data!</p>\n<p>The key for me was to figure out how to reduce that count. Averaging can work well (and there are other methods as well). For example - here's a path with 3 waypoints with all of the acceleration data, and then that data averaged into 100 points: <a href=\"https://imgur.com/a/sTZsp3i\" target=\"_blank\">https://imgur.com/a/sTZsp3i</a> (kaggle isn't allowing file uploads right now, so I had to use imgur).</p>\n<p>The key to figure out the right level of data reduction is to do EDA (exploratory data analysis), and a lot of graphing; and then a way to feed that into my model.</p>\n<p><strong>2. The \"z\" axis is the most important</strong></p>\n<p>The project github link (<a href=\"https://github.com/location-competition/indoor-location-competition-20\" target=\"_blank\">https://github.com/location-competition/indoor-location-competition-20</a>) has a link to the Android sensor data page (<a href=\"https://developer.android.com/guide/topics/sensors/sensors_overview\" target=\"_blank\">https://developer.android.com/guide/topics/sensors/sensors_overview</a>) which has an image describing the sensor coordinate system. When you combine that with the description of how the data was collected (the phone was held flat in front of the user), it's clear that the \"z\" axis will be the most important for the sensor values (based on how people move and rotate if they are holding a phone flat in front of them).</p>\n<p>In my own tests, I was able to see about 80% of the value by just using the z axis - so if you're stuck with too much data, start there.</p>\n<p><strong>3. RNNs are difficult to train</strong></p>\n<p>I had a lot of trouble getting an RNN or LSTM to train on the sensor data - but there's more than one way to handle a time series! You can start with just a regular feed forward neural net, a CNN (my favorite), a transformer, or even another ML method like xgboost if you engineer features first.</p>\n<p>Don't get stuck thinking you need a recurrent network just because the popular notebooks use one - you can handle this data in a number of different ways to get value from it.</p>\n<p>You may also think that you need to understand Kalman filters (or similar) before you can use sensor data - but not so! You can start using sensor data in neural nets (for example) without understanding it at all :)</p>\n<p><strong>Overall</strong></p>\n<p>I would start by graphing a bunch of the sensor data yourself - you'll start to see patterns in the paths, around the waypoints, etc, and that should guide you with how to use it.</p>\n<p>I hope this helps some!</p>\n<p>I couldn't get past ~4.5m with just wifi, so hopefully adding some sensor data will help you too :)</p>",
      "rawMarkdown": "If you're having trouble getting past the 4-5m mark and haven't looked at the sensor data yet, here are a few tricks to get you started:\n\n**1. There's too much sensor data**\n\nThe longest test path has 36,577 sensor points, each with 12 values - that's a lot of data!\n\nThe key for me was to figure out how to reduce that count. Averaging can work well (and there are other methods as well). For example - here's a path with 3 waypoints with all of the acceleration data, and then that data averaged into 100 points: https://imgur.com/a/sTZsp3i (kaggle isn't allowing file uploads right now, so I had to use imgur).\n\nThe key to figure out the right level of data reduction is to do EDA (exploratory data analysis), and a lot of graphing; and then a way to feed that into my model.\n\n**2. The \"z\" axis is the most important**\n\nThe project github link (https://github.com/location-competition/indoor-location-competition-20) has a link to the Android sensor data page (https://developer.android.com/guide/topics/sensors/sensors_overview) which has an image describing the sensor coordinate system. When you combine that with the description of how the data was collected (the phone was held flat in front of the user), it's clear that the \"z\" axis will be the most important for the sensor values (based on how people move and rotate if they are holding a phone flat in front of them).\n\nIn my own tests, I was able to see about 80% of the value by just using the z axis - so if you're stuck with too much data, start there.\n\n**3. RNNs are difficult to train**\n\nI had a lot of trouble getting an RNN or LSTM to train on the sensor data - but there's more than one way to handle a time series! You can start with just a regular feed forward neural net, a CNN (my favorite), a transformer, or even another ML method like xgboost if you engineer features first.\n\nDon't get stuck thinking you need a recurrent network just because the popular notebooks use one - you can handle this data in a number of different ways to get value from it.\n\nYou may also think that you need to understand Kalman filters (or similar) before you can use sensor data - but not so! You can start using sensor data in neural nets (for example) without understanding it at all :)\n\n**Overall**\n\nI would start by graphing a bunch of the sensor data yourself - you'll start to see patterns in the paths, around the waypoints, etc, and that should guide you with how to use it.\n\nI hope this helps some!\n\nI couldn't get past ~4.5m with just wifi, so hopefully adding some sensor data will help you too :)",
      "votes": null
    },
    {
      "id": "1291791",
      "postDate": "05/03/2021 10:50:52",
      "content": "<p>Thanks for sharing, this is valuable advice, myself stuck at 4-5 range :(</p>",
      "rawMarkdown": "Thanks for sharing, this is valuable advice, myself stuck at 4-5 range :(",
      "votes": null
    },
    {
      "id": "1292935",
      "postDate": "05/04/2021 12:30:00",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chris62\" target=\"_blank\">@chris62</a>, thank you for your generous sharing. May I ask when you got stuck with ~4.5m using wifi, was it the error before or after using post-processing?</p>",
      "rawMarkdown": "Hi @chris62, thank you for your generous sharing. May I ask when you got stuck with ~4.5m using wifi, was it the error before or after using post-processing?",
      "votes": null
    },
    {
      "id": "1292950",
      "postDate": "05/04/2021 12:43:04",
      "content": "<p>That was after post processing - but maybe my post processing wasn't all that good.</p>\n<p>If I went back to just wifi with some of the post processing I'm using now, then perhaps I could get down to 4m (not sure though), but I still think it would be difficult to go past that.  I'm not sure though! maybe someone is doing it already with some neat tricks - I'm excited to hear about it at the end of the competition if so :)</p>",
      "rawMarkdown": "That was after post processing - but maybe my post processing wasn't all that good.\n\nIf I went back to just wifi with some of the post processing I'm using now, then perhaps I could get down to 4m (not sure though), but I still think it would be difficult to go past that.  I'm not sure though! maybe someone is doing it already with some neat tricks - I'm excited to hear about it at the end of the competition if so :)",
      "votes": null
    },
    {
      "id": "1293024",
      "postDate": "05/04/2021 13:37:04",
      "content": "<p>thank you for your clarification. We also thought about using IMU along with Wifi readings in the training process. We actually applied some quite complicated techniques, but so far, they failed or brought marginal effects. I am also excited to see your solution as well as other top teams' solutions at the end of this competition :D</p>",
      "rawMarkdown": "thank you for your clarification. We also thought about using IMU along with Wifi readings in the training process. We actually applied some quite complicated techniques, but so far, they failed or brought marginal effects. I am also excited to see your solution as well as other top teams' solutions at the end of this competition :D",
      "votes": null
    },
    {
      "id": "1293035",
      "postDate": "05/04/2021 13:47:24",
      "content": "<p>I also started by trying very complicated techniques, but as I mentioned in the post, the real trick for me was to figure out how to reduce the amount of data and the complexity - so I guess I'd say, start with just the simplest thing you can think of, and expand from there. Good luck!</p>",
      "rawMarkdown": "I also started by trying very complicated techniques, but as I mentioned in the post, the real trick for me was to figure out how to reduce the amount of data and the complexity - so I guess I'd say, start with just the simplest thing you can think of, and expand from there. Good luck!",
      "votes": null
    },
    {
      "id": "1293043",
      "postDate": "05/04/2021 13:57:22",
      "content": "<p>thank you very much for your sharing :D, we appreciate it.</p>",
      "rawMarkdown": "thank you very much for your sharing :D, we appreciate it.",
      "votes": null
    },
    {
      "id": "1293515",
      "postDate": "05/05/2021 00:51:28",
      "content": "<p>Thank you for sharing. You seem to know lots of things we don't know 😲</p>",
      "rawMarkdown": "Thank you for sharing. You seem to know lots of things we don't know 😲",
      "votes": null
    },
    {
      "id": "1402402",
      "postDate": "07/28/2021 06:58:20",
      "content": "<p>Hi~ Do i need to use kalman filter before feed data into LSTM? or just feed the raw data?</p>",
      "rawMarkdown": "Hi~ Do i need to use kalman filter before feed data into LSTM? or just feed the raw data?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1291791,
      "author_name": "saurabhbagchi",
      "author_url": "",
      "post_date": "05/03/2021 10:50:52",
      "content": "<p>Thanks for sharing, this is valuable advice, myself stuck at 4-5 range :(</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1292935,
      "author_name": "shinomoriaoshi",
      "author_url": "",
      "post_date": "05/04/2021 12:30:00",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chris62\" target=\"_blank\">@chris62</a>, thank you for your generous sharing. May I ask when you got stuck with ~4.5m using wifi, was it the error before or after using post-processing?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1292950,
          "author_name": "chris62",
          "author_url": "",
          "post_date": "05/04/2021 12:43:04",
          "content": "<p>That was after post processing - but maybe my post processing wasn't all that good.</p>\n<p>If I went back to just wifi with some of the post processing I'm using now, then perhaps I could get down to 4m (not sure though), but I still think it would be difficult to go past that.  I'm not sure though! maybe someone is doing it already with some neat tricks - I'm excited to hear about it at the end of the competition if so :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1293024,
          "author_name": "shinomoriaoshi",
          "author_url": "",
          "post_date": "05/04/2021 13:37:04",
          "content": "<p>thank you for your clarification. We also thought about using IMU along with Wifi readings in the training process. We actually applied some quite complicated techniques, but so far, they failed or brought marginal effects. I am also excited to see your solution as well as other top teams' solutions at the end of this competition :D</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1293035,
          "author_name": "chris62",
          "author_url": "",
          "post_date": "05/04/2021 13:47:24",
          "content": "<p>I also started by trying very complicated techniques, but as I mentioned in the post, the real trick for me was to figure out how to reduce the amount of data and the complexity - so I guess I'd say, start with just the simplest thing you can think of, and expand from there. Good luck!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1293043,
          "author_name": "shinomoriaoshi",
          "author_url": "",
          "post_date": "05/04/2021 13:57:22",
          "content": "<p>thank you very much for your sharing :D, we appreciate it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1293515,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "05/05/2021 00:51:28",
      "content": "<p>Thank you for sharing. You seem to know lots of things we don't know 😲</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1402402,
      "author_name": "edwintsang",
      "author_url": "",
      "post_date": "07/28/2021 06:58:20",
      "content": "<p>Hi~ Do i need to use kalman filter before feed data into LSTM? or just feed the raw data?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1291248": "If you're having trouble getting past the 4-5m mark and haven't looked at the sensor data yet, here are a few tricks to get you started:\n\n**1. There's too much sensor data**\n\nThe longest test path has 36,577 sensor points, each with 12 values - that's a lot of data!\n\nThe key for me was to figure out how to reduce that count. Averaging can work well (and there are other methods as well). For example - here's a path with 3 waypoints with all of the acceleration data, and then that data averaged into 100 points: https://imgur.com/a/sTZsp3i (kaggle isn't allowing file uploads right now, so I had to use imgur).\n\nThe key to figure out the right level of data reduction is to do EDA (exploratory data analysis), and a lot of graphing; and then a way to feed that into my model.\n\n**2. The \"z\" axis is the most important**\n\nThe project github link (https://github.com/location-competition/indoor-location-competition-20) has a link to the Android sensor data page (https://developer.android.com/guide/topics/sensors/sensors_overview) which has an image describing the sensor coordinate system. When you combine that with the description of how the data was collected (the phone was held flat in front of the user), it's clear that the \"z\" axis will be the most important for the sensor values (based on how people move and rotate if they are holding a phone flat in front of them).\n\nIn my own tests, I was able to see about 80% of the value by just using the z axis - so if you're stuck with too much data, start there.\n\n**3. RNNs are difficult to train**\n\nI had a lot of trouble getting an RNN or LSTM to train on the sensor data - but there's more than one way to handle a time series! You can start with just a regular feed forward neural net, a CNN (my favorite), a transformer, or even another ML method like xgboost if you engineer features first.\n\nDon't get stuck thinking you need a recurrent network just because the popular notebooks use one - you can handle this data in a number of different ways to get value from it.\n\nYou may also think that you need to understand Kalman filters (or similar) before you can use sensor data - but not so! You can start using sensor data in neural nets (for example) without understanding it at all :)\n\n**Overall**\n\nI would start by graphing a bunch of the sensor data yourself - you'll start to see patterns in the paths, around the waypoints, etc, and that should guide you with how to use it.\n\nI hope this helps some!\n\nI couldn't get past ~4.5m with just wifi, so hopefully adding some sensor data will help you too :)",
    "1291791": "Thanks for sharing, this is valuable advice, myself stuck at 4-5 range :(",
    "1292935": "Hi @chris62, thank you for your generous sharing. May I ask when you got stuck with ~4.5m using wifi, was it the error before or after using post-processing?",
    "1292950": "That was after post processing - but maybe my post processing wasn't all that good.\n\nIf I went back to just wifi with some of the post processing I'm using now, then perhaps I could get down to 4m (not sure though), but I still think it would be difficult to go past that.  I'm not sure though! maybe someone is doing it already with some neat tricks - I'm excited to hear about it at the end of the competition if so :)",
    "1293024": "thank you for your clarification. We also thought about using IMU along with Wifi readings in the training process. We actually applied some quite complicated techniques, but so far, they failed or brought marginal effects. I am also excited to see your solution as well as other top teams' solutions at the end of this competition :D",
    "1293035": "I also started by trying very complicated techniques, but as I mentioned in the post, the real trick for me was to figure out how to reduce the amount of data and the complexity - so I guess I'd say, start with just the simplest thing you can think of, and expand from there. Good luck!",
    "1293043": "thank you very much for your sharing :D, we appreciate it.",
    "1293515": "Thank you for sharing. You seem to know lots of things we don't know 😲",
    "1402402": "Hi~ Do i need to use kalman filter before feed data into LSTM? or just feed the raw data?"
  },
  "source": "meta"
}