{
  "id": 215381,
  "title": "Reading the trace files with pandas",
  "url": "/competitions/indoor-location-navigation/discussion/215381",
  "author_name": "",
  "post_date": "2021-01-29T15:27:28.603169800Z",
  "votes": 45,
  "comment_count": 8,
  "views": 0,
  "content": "<p>The trace files seem to be complicated at first glance. Here I show how to read a trace file with pandas. I hope you find it useful.</p>\n<pre><code>import pandas as pd\n\npath_trace_file = '../input/indoor-location-navigation/train/5a0546857ecc773753327266/F4/5d11dbf8ffe23f0008604f53.txt'\n\nnames = ['Time', 'Type'] + ['col'+str(x) for x in range(1,9)]\ndf = pd.read_csv(path_trace_file, sep='\\t', comment='#', header=None, names=names)\n</code></pre>\n<p>Note that there are at most 8 columns in addition to Time and Type columns. After reading the file, don't forget to assign the appropriate dtypes.</p>\n<p>A demonstration of how to fully parse a single path trace file is available <a href=\"https://www.kaggle.com/tolgadincer/iln-quick-overview\" target=\"_blank\"><strong>here</strong>.</a></p>",
  "messages": [
    {
      "id": "1176334",
      "postDate": "01/29/2021 15:27:28",
      "content": "<p>The trace files seem to be complicated at first glance. Here I show how to read a trace file with pandas. I hope you find it useful.</p>\n<pre><code>import pandas as pd\n\npath_trace_file = '../input/indoor-location-navigation/train/5a0546857ecc773753327266/F4/5d11dbf8ffe23f0008604f53.txt'\n\nnames = ['Time', 'Type'] + ['col'+str(x) for x in range(1,9)]\ndf = pd.read_csv(path_trace_file, sep='\\t', comment='#', header=None, names=names)\n</code></pre>\n<p>Note that there are at most 8 columns in addition to Time and Type columns. After reading the file, don't forget to assign the appropriate dtypes.</p>\n<p>A demonstration of how to fully parse a single path trace file is available <a href=\"https://www.kaggle.com/tolgadincer/iln-quick-overview\" target=\"_blank\"><strong>here</strong>.</a></p>",
      "rawMarkdown": "The trace files seem to be complicated at first glance. Here I show how to read a trace file with pandas. I hope you find it useful.\n\n```python\nimport pandas as pd\n\npath_trace_file = '../input/indoor-location-navigation/train/5a0546857ecc773753327266/F4/5d11dbf8ffe23f0008604f53.txt'\n\nnames = ['Time', 'Type'] + ['col'+str(x) for x in range(1,9)]\ndf = pd.read_csv(path_trace_file, sep='\\t', comment='#', header=None, names=names)\n```\n\nNote that there are at most 8 columns in addition to Time and Type columns. After reading the file, don't forget to assign the appropriate dtypes.\n\nA demonstration of how to fully parse a single path trace file is available [**here**.](https://www.kaggle.com/tolgadincer/iln-quick-overview)",
      "votes": null
    },
    {
      "id": "1176389",
      "postDate": "01/29/2021 15:52:21",
      "content": "<p>Very helpful, thank you!</p>",
      "rawMarkdown": "Very helpful, thank you!",
      "votes": null
    },
    {
      "id": "1178472",
      "postDate": "01/30/2021 22:35:02",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/tolgadincer\" target=\"_blank\">@tolgadincer</a> !</p>\n<p>Your script is very useful, thank you for sharing! </p>\n<p>I want to ask you about the reason why the data contains at most 8 columns? Maybe it was said in the documentation or you check it in the data?</p>\n<p>P.S. Sorry if I missed this information in the docs.</p>",
      "rawMarkdown": "Hello @tolgadincer !\n\nYour script is very useful, thank you for sharing! \n\nI want to ask you about the reason why the data contains at most 8 columns? Maybe it was said in the documentation or you check it in the data?\n\nP.S. Sorry if I missed this information in the docs.",
      "votes": null
    },
    {
      "id": "1178551",
      "postDate": "01/31/2021 01:05:46",
      "content": "<p>I'm glad you find it useful!</p>\n<p>Just to make it clear, my statement is that there are 8 columns in addition to the TIME and TYPE columns. So, there are 10 columns in total. I parsed several different path trace files and pandas did not find any data after the 10th column.</p>\n<p>However, there are reports of corrupted files. Those files need special care when parsing (e.g. see the file in <a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/215308\" target=\"_blank\">this post</a>). </p>",
      "rawMarkdown": "I'm glad you find it useful!\n\nJust to make it clear, my statement is that there are 8 columns in addition to the TIME and TYPE columns. So, there are 10 columns in total. I parsed several different path trace files and pandas did not find any data after the 10th column.\n\nHowever, there are reports of corrupted files. Those files need special care when parsing (e.g. see the file in [this post](https://www.kaggle.com/c/indoor-location-navigation/discussion/215308)).",
      "votes": null
    },
    {
      "id": "1179369",
      "postDate": "01/31/2021 14:23:35",
      "content": "<p>Thanks for the answer!</p>\n<p>Now it’s clear to me.</p>",
      "rawMarkdown": "Thanks for the answer!\n\nNow it’s clear to me.",
      "votes": null
    },
    {
      "id": "1183566",
      "postDate": "02/03/2021 05:10:45",
      "content": "<p><a href=\"https://www.kaggle.com/tosinabase\" target=\"_blank\">@tosinabase</a> and <a href=\"https://www.kaggle.com/tolgadincer\" target=\"_blank\">@tolgadincer</a> you can see from the Data section of the competition there is a github link given for some basic explanation and processing of data, there you can find that there are atmost 10 columns.<br>\nFor your reference here is the github link: <a href=\"https://github.com/location-competition/indoor-location-competition-20\" target=\"_blank\">https://github.com/location-competition/indoor-location-competition-20</a></p>",
      "rawMarkdown": "tosinabase and @tolgadincer you can see from the Data section of the competition there is a github link given for some basic explanation and processing of data, there you can find that there are atmost 10 columns.\nFor your reference here is the github link: https://github.com/location-competition/indoor-location-competition-20",
      "votes": null
    },
    {
      "id": "1204472",
      "postDate": "02/16/2021 06:48:27",
      "content": "<p>Awesome this makes it so much easy to enter the competition. As a novice, I have been trying to make sense of the data but did not have much luck. This surely helps a lot <a href=\"https://www.kaggle.com/tolgadincer\" target=\"_blank\">@tolgadincer</a> thanks. </p>",
      "rawMarkdown": "Awesome this makes it so much easy to enter the competition. As a novice, I have been trying to make sense of the data but did not have much luck. This surely helps a lot @tolgadincer thanks.",
      "votes": null
    },
    {
      "id": "1228564",
      "postDate": "03/06/2021 14:53:43",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/tolgadincer\" target=\"_blank\">@tolgadincer</a> for sharing these scripts. I am just starting on this competition and your scripts are very helpful. </p>",
      "rawMarkdown": "Thanks @tolgadincer for sharing these scripts. I am just starting on this competition and your scripts are very helpful.",
      "votes": null
    },
    {
      "id": "1289668",
      "postDate": "05/01/2021 08:49:02",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/tolgadincer\" target=\"_blank\">@tolgadincer</a> !<br>\nI have just joined. This scripts very helpful !!</p>",
      "rawMarkdown": "Thanks @tolgadincer !\nI have just joined. This scripts very helpful !!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1176389,
      "author_name": "jonas0",
      "author_url": "",
      "post_date": "01/29/2021 15:52:21",
      "content": "<p>Very helpful, thank you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1178472,
      "author_name": "tosinabase",
      "author_url": "",
      "post_date": "01/30/2021 22:35:02",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/tolgadincer\" target=\"_blank\">@tolgadincer</a> !</p>\n<p>Your script is very useful, thank you for sharing! </p>\n<p>I want to ask you about the reason why the data contains at most 8 columns? Maybe it was said in the documentation or you check it in the data?</p>\n<p>P.S. Sorry if I missed this information in the docs.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1178551,
          "author_name": "tolgadincer",
          "author_url": "",
          "post_date": "01/31/2021 01:05:46",
          "content": "<p>I'm glad you find it useful!</p>\n<p>Just to make it clear, my statement is that there are 8 columns in addition to the TIME and TYPE columns. So, there are 10 columns in total. I parsed several different path trace files and pandas did not find any data after the 10th column.</p>\n<p>However, there are reports of corrupted files. Those files need special care when parsing (e.g. see the file in <a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/215308\" target=\"_blank\">this post</a>). </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1179369,
          "author_name": "tosinabase",
          "author_url": "",
          "post_date": "01/31/2021 14:23:35",
          "content": "<p>Thanks for the answer!</p>\n<p>Now it’s clear to me.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1183566,
          "author_name": "smritisingh1997",
          "author_url": "",
          "post_date": "02/03/2021 05:10:45",
          "content": "<p><a href=\"https://www.kaggle.com/tosinabase\" target=\"_blank\">@tosinabase</a> and <a href=\"https://www.kaggle.com/tolgadincer\" target=\"_blank\">@tolgadincer</a> you can see from the Data section of the competition there is a github link given for some basic explanation and processing of data, there you can find that there are atmost 10 columns.<br>\nFor your reference here is the github link: <a href=\"https://github.com/location-competition/indoor-location-competition-20\" target=\"_blank\">https://github.com/location-competition/indoor-location-competition-20</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1204472,
      "author_name": "kanishkpatel",
      "author_url": "",
      "post_date": "02/16/2021 06:48:27",
      "content": "<p>Awesome this makes it so much easy to enter the competition. As a novice, I have been trying to make sense of the data but did not have much luck. This surely helps a lot <a href=\"https://www.kaggle.com/tolgadincer\" target=\"_blank\">@tolgadincer</a> thanks. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1228564,
      "author_name": "suryajrrafl",
      "author_url": "",
      "post_date": "03/06/2021 14:53:43",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/tolgadincer\" target=\"_blank\">@tolgadincer</a> for sharing these scripts. I am just starting on this competition and your scripts are very helpful. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1289668,
      "author_name": "ryotak12",
      "author_url": "",
      "post_date": "05/01/2021 08:49:02",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/tolgadincer\" target=\"_blank\">@tolgadincer</a> !<br>\nI have just joined. This scripts very helpful !!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1176334": "The trace files seem to be complicated at first glance. Here I show how to read a trace file with pandas. I hope you find it useful.\n\n```python\nimport pandas as pd\n\npath_trace_file = '../input/indoor-location-navigation/train/5a0546857ecc773753327266/F4/5d11dbf8ffe23f0008604f53.txt'\n\nnames = ['Time', 'Type'] + ['col'+str(x) for x in range(1,9)]\ndf = pd.read_csv(path_trace_file, sep='\\t', comment='#', header=None, names=names)\n```\n\nNote that there are at most 8 columns in addition to Time and Type columns. After reading the file, don't forget to assign the appropriate dtypes.\n\nA demonstration of how to fully parse a single path trace file is available [**here**.](https://www.kaggle.com/tolgadincer/iln-quick-overview)",
    "1176389": "Very helpful, thank you!",
    "1178472": "Hello @tolgadincer !\n\nYour script is very useful, thank you for sharing! \n\nI want to ask you about the reason why the data contains at most 8 columns? Maybe it was said in the documentation or you check it in the data?\n\nP.S. Sorry if I missed this information in the docs.",
    "1178551": "I'm glad you find it useful!\n\nJust to make it clear, my statement is that there are 8 columns in addition to the TIME and TYPE columns. So, there are 10 columns in total. I parsed several different path trace files and pandas did not find any data after the 10th column.\n\nHowever, there are reports of corrupted files. Those files need special care when parsing (e.g. see the file in [this post](https://www.kaggle.com/c/indoor-location-navigation/discussion/215308)).",
    "1179369": "Thanks for the answer!\n\nNow it’s clear to me.",
    "1183566": "tosinabase and @tolgadincer you can see from the Data section of the competition there is a github link given for some basic explanation and processing of data, there you can find that there are atmost 10 columns.\nFor your reference here is the github link: https://github.com/location-competition/indoor-location-competition-20",
    "1204472": "Awesome this makes it so much easy to enter the competition. As a novice, I have been trying to make sense of the data but did not have much luck. This surely helps a lot @tolgadincer thanks.",
    "1228564": "Thanks @tolgadincer for sharing these scripts. I am just starting on this competition and your scripts are very helpful.",
    "1289668": "Thanks @tolgadincer !\nI have just joined. This scripts very helpful !!"
  },
  "source": "meta"
}