{
  "id": 215973,
  "title": "Some of the txt data are broken.",
  "url": "/competitions/indoor-location-navigation/discussion/215973",
  "author_name": "kenmatsu4",
  "post_date": "2021-02-01T04:48:42.105000",
  "votes": 25,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I found some of the txt data are broken. <br>\nMultiple <code>TYPE_</code> data exist in one line. For example, TYPE_BEACON and TYPE_ACCELEROMETER  are on the one line becase of lack of return code indicated at the below image.<br>\nPlease pay attention, all!</p>\n<p><img src=\"https://user-images.githubusercontent.com/7313648/106415914-9e13cd00-6493-11eb-8a84-1c819018c47b.png\" alt=\"\"></p>\n<p>sample txt path:<br>\n<code>../input/indoor-location-navigation/train/5cd56ba1e2acfd2d33b603af/B2/5cf75ac5e36a480008125f9d.txt</code></p>\n<p>Solution:<br>\nThis problem is handled on my feature store class. Please use if you like.<br>\n<a href=\"https://www.kaggle.com/kenmatsu4/feature-store-for-indoor-location-navigation\" target=\"_blank\">https://www.kaggle.com/kenmatsu4/feature-store-for-indoor-location-navigation</a></p>",
  "messages": [
    {
      "id": 1180087,
      "postDate": "2021-02-01T04:48:42.107Z",
      "content": "<p>I found some of the txt data are broken. <br>\nMultiple <code>TYPE_</code> data exist in one line. For example, TYPE_BEACON and TYPE_ACCELEROMETER  are on the one line becase of lack of return code indicated at the below image.<br>\nPlease pay attention, all!</p>\n<p><img src=\"https://user-images.githubusercontent.com/7313648/106415914-9e13cd00-6493-11eb-8a84-1c819018c47b.png\" alt=\"\"></p>\n<p>sample txt path:<br>\n<code>../input/indoor-location-navigation/train/5cd56ba1e2acfd2d33b603af/B2/5cf75ac5e36a480008125f9d.txt</code></p>\n<p>Solution:<br>\nThis problem is handled on my feature store class. Please use if you like.<br>\n<a href=\"https://www.kaggle.com/kenmatsu4/feature-store-for-indoor-location-navigation\" target=\"_blank\">https://www.kaggle.com/kenmatsu4/feature-store-for-indoor-location-navigation</a></p>",
      "rawMarkdown": "I found some of the txt data are broken. \nMultiple `TYPE_` data exist in one line. For example, TYPE_BEACON and TYPE_ACCELEROMETER  are on the one line becase of lack of return code indicated at the below image.\nPlease pay attention, all!\n\n![](https://user-images.githubusercontent.com/7313648/106415914-9e13cd00-6493-11eb-8a84-1c819018c47b.png)\n\nsample txt path:\n`../input/indoor-location-navigation/train/5cd56ba1e2acfd2d33b603af/B2/5cf75ac5e36a480008125f9d.txt`\n\nSolution:\nThis problem is handled on my feature store class. Please use if you like.\nhttps://www.kaggle.com/kenmatsu4/feature-store-for-indoor-location-navigation",
      "votes": 25
    },
    {
      "id": 1180231,
      "postDate": "2021-02-01T07:04:14.257Z",
      "content": "<p>Thank you kenmatsu for great notebook😍<br>\nThe key idea seems here.</p>\n<pre><code>    def multi_line_spliter(self, s):\n        matches = re.finditer(\"TYPE_\", s)\n        matches_positions = [match.start() for match in matches]\n        split_idx = [0] + [matches_positions[i]-14 for i in range(1, len(matches_positions))] + [len(s)]\n        return [s[split_idx[i]:split_idx[i+1]] for i in range(len(split_idx)-1)]\n</code></pre>\n<p>In my understanding, 14 means <code>the length of the timestamp + 1</code>. It's very reasonable!<br>\nI hope it helps people who want to use only the part of kenmatsu's code🙏 </p>",
      "rawMarkdown": "Thank you kenmatsu for great notebook😍\nThe key idea seems here.\n```\n    def multi_line_spliter(self, s):\n        matches = re.finditer(\"TYPE_\", s)\n        matches_positions = [match.start() for match in matches]\n        split_idx = [0] + [matches_positions[i]-14 for i in range(1, len(matches_positions))] + [len(s)]\n        return [s[split_idx[i]:split_idx[i+1]] for i in range(len(split_idx)-1)]\n```\nIn my understanding, 14 means `the length of the timestamp + 1`. It's very reasonable!\nI hope it helps people who want to use only the part of kenmatsu's code🙏 ",
      "votes": 7,
      "replies": [
        {
          "id": 1237517,
          "postDate": "2021-03-14T08:14:41.667Z",
          "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a>, can you tell me where is he saving all the processed text files, as far as I understood from the notebook, all the features are individually there in the 'feature' variable and we can access them by looping, but is there some way to combine them and save them so that these can be used later for further preprocessing?</p>",
          "rawMarkdown": "@mamasinkgs, can you tell me where is he saving all the processed text files, as far as I understood from the notebook, all the features are individually there in the 'feature' variable and we can access them by looping, but is there some way to combine them and save them so that these can be used later for further preprocessing?"
        }
      ]
    },
    {
      "id": 1197490,
      "postDate": "2021-02-12T07:20:08.733Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kenmatsu4\" target=\"_blank\">@kenmatsu4</a>, does the solution provided by you also solves the following issues which are present in the files:</p>\n<ol>\n<li>Number of footer row is not fixed.</li>\n<li>A file may have no footer.</li>\n<li>About header format, field name and value is not always separated by \":\", sometimes by \"\\t\".</li>\n</ol>",
      "rawMarkdown": "Hi @kenmatsu4, does the solution provided by you also solves the following issues which are present in the files:\n1.  Number of footer row is not fixed.\n2. A file may have no footer.\n3. About header format, field name and value is not always separated by \":\", sometimes by \"\\t\"."
    },
    {
      "id": 1244152,
      "postDate": "2021-03-18T18:11:36.157Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1180231,
      "author_name": "mamas",
      "author_url": "",
      "post_date": "2021-02-01T07:04:14.257000",
      "content": "<p>Thank you kenmatsu for great notebook😍<br>\nThe key idea seems here.</p>\n<pre><code>    def multi_line_spliter(self, s):\n        matches = re.finditer(\"TYPE_\", s)\n        matches_positions = [match.start() for match in matches]\n        split_idx = [0] + [matches_positions[i]-14 for i in range(1, len(matches_positions))] + [len(s)]\n        return [s[split_idx[i]:split_idx[i+1]] for i in range(len(split_idx)-1)]\n</code></pre>\n<p>In my understanding, 14 means <code>the length of the timestamp + 1</code>. It's very reasonable!<br>\nI hope it helps people who want to use only the part of kenmatsu's code🙏 </p>",
      "votes": 7,
      "replies": [
        {
          "id": 1237517,
          "author_name": "Smriti ",
          "author_url": "",
          "post_date": "2021-03-14T08:14:41.667000",
          "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a>, can you tell me where is he saving all the processed text files, as far as I understood from the notebook, all the features are individually there in the 'feature' variable and we can access them by looping, but is there some way to combine them and save them so that these can be used later for further preprocessing?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1197490,
      "author_name": "Smriti ",
      "author_url": "",
      "post_date": "2021-02-12T07:20:08.733000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kenmatsu4\" target=\"_blank\">@kenmatsu4</a>, does the solution provided by you also solves the following issues which are present in the files:</p>\n<ol>\n<li>Number of footer row is not fixed.</li>\n<li>A file may have no footer.</li>\n<li>About header format, field name and value is not always separated by \":\", sometimes by \"\\t\".</li>\n</ol>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1244152,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-18T18:11:36.157000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1180087": "I found some of the txt data are broken. \nMultiple `TYPE_` data exist in one line. For example, TYPE_BEACON and TYPE_ACCELEROMETER  are on the one line becase of lack of return code indicated at the below image.\nPlease pay attention, all!\n\n![](https://user-images.githubusercontent.com/7313648/106415914-9e13cd00-6493-11eb-8a84-1c819018c47b.png)\n\nsample txt path:\n`../input/indoor-location-navigation/train/5cd56ba1e2acfd2d33b603af/B2/5cf75ac5e36a480008125f9d.txt`\n\nSolution:\nThis problem is handled on my feature store class. Please use if you like.\nhttps://www.kaggle.com/kenmatsu4/feature-store-for-indoor-location-navigation",
    "1180231": "Thank you kenmatsu for great notebook😍\nThe key idea seems here.\n```\n    def multi_line_spliter(self, s):\n        matches = re.finditer(\"TYPE_\", s)\n        matches_positions = [match.start() for match in matches]\n        split_idx = [0] + [matches_positions[i]-14 for i in range(1, len(matches_positions))] + [len(s)]\n        return [s[split_idx[i]:split_idx[i+1]] for i in range(len(split_idx)-1)]\n```\nIn my understanding, 14 means `the length of the timestamp + 1`. It's very reasonable!\nI hope it helps people who want to use only the part of kenmatsu's code🙏 ",
    "1197490": "Hi @kenmatsu4, does the solution provided by you also solves the following issues which are present in the files:\n1.  Number of footer row is not fixed.\n2. A file may have no footer.\n3. About header format, field name and value is not always separated by \":\", sometimes by \"\\t\".",
    "1244152": ""
  }
}