{
  "id": 225353,
  "title": "fix for host's io_f.py to handle missed new lines ",
  "url": "/competitions/indoor-location-navigation/discussion/225353",
  "author_name": "Evgeniya",
  "post_date": "2021-03-11T21:30:06.334000",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi, <br>\nas you know sometimes \\n are missed in data files. To use io_f.read_data_file (from the repository provided) in your code, the following patch could be useful, the whole file is available at <a href=\"url\" target=\"_blank\">https://www.kaggle.com/egm108/fix-for-io-f-py</a></p>\n<pre><code> import numpy as np\n-\n+import re\n\n @dataclass\n class ReadData:\n@@ -31,8 +31,20 @@ def read_data_file(data_filename):\n\n     with open(data_filename, 'r', encoding='utf-8') as file:\n         lines = file.readlines()\n-\n-    for line_data in lines:\n+\n+    new_list = []\n+    for line in lines:\n+        if (\"TYPE_\" in line[30:]):\n+            a = [m.start(0) for m in re.finditer(r'\\d{13}\\tTYPE', line[30:])]\n+            new_list.append(line[0:30 + a[0]] )\n+            for i in range(len(a)-1):\n+                new_list.append(line[30:][a[i]:a[i+1]])\n+\n+            new_list.append(line[30:][a[-1]:])\n+        else:\n+            new_list.append(line)\n+\n+    for line_data in new_list:\n         line_data = line_data.strip()\n         if not line_data or line_data[0] == '#':\n:\n</code></pre>\n<p>Thanks</p>",
  "messages": [
    {
      "id": 1235135,
      "postDate": "2021-03-11T21:30:06.333Z",
      "content": "<p>Hi, <br>\nas you know sometimes \\n are missed in data files. To use io_f.read_data_file (from the repository provided) in your code, the following patch could be useful, the whole file is available at <a href=\"url\" target=\"_blank\">https://www.kaggle.com/egm108/fix-for-io-f-py</a></p>\n<pre><code> import numpy as np\n-\n+import re\n\n @dataclass\n class ReadData:\n@@ -31,8 +31,20 @@ def read_data_file(data_filename):\n\n     with open(data_filename, 'r', encoding='utf-8') as file:\n         lines = file.readlines()\n-\n-    for line_data in lines:\n+\n+    new_list = []\n+    for line in lines:\n+        if (\"TYPE_\" in line[30:]):\n+            a = [m.start(0) for m in re.finditer(r'\\d{13}\\tTYPE', line[30:])]\n+            new_list.append(line[0:30 + a[0]] )\n+            for i in range(len(a)-1):\n+                new_list.append(line[30:][a[i]:a[i+1]])\n+\n+            new_list.append(line[30:][a[-1]:])\n+        else:\n+            new_list.append(line)\n+\n+    for line_data in new_list:\n         line_data = line_data.strip()\n         if not line_data or line_data[0] == '#':\n:\n</code></pre>\n<p>Thanks</p>",
      "rawMarkdown": "Hi, \nas you know sometimes \\n are missed in data files. To use io_f.read_data_file (from the repository provided) in your code, the following patch could be useful, the whole file is available at [https://www.kaggle.com/egm108/fix-for-io-f-py](url)\n\n```\n import numpy as np\n-\n+import re\n\n @dataclass\n class ReadData:\n@@ -31,8 +31,20 @@ def read_data_file(data_filename):\n\n     with open(data_filename, 'r', encoding='utf-8') as file:\n         lines = file.readlines()\n-\n-    for line_data in lines:\n+\n+    new_list = []\n+    for line in lines:\n+        if (\"TYPE_\" in line[30:]):\n+            a = [m.start(0) for m in re.finditer(r'\\d{13}\\tTYPE', line[30:])]\n+            new_list.append(line[0:30 + a[0]] )\n+            for i in range(len(a)-1):\n+                new_list.append(line[30:][a[i]:a[i+1]])\n+\n+            new_list.append(line[30:][a[-1]:])\n+        else:\n+            new_list.append(line)\n+\n+    for line_data in new_list:\n         line_data = line_data.strip()\n         if not line_data or line_data[0] == '#':\n:\n\n```\n\nThanks",
      "votes": 2
    },
    {
      "id": 1256663,
      "postDate": "2021-03-30T06:21:14.207Z",
      "content": "<p>Thanks for sharing! I did not get how do you deal with the timestamps. They are normally added before the wort \"TYPE\". As for example here: </p>\n<p>\"1559699290975    TYPE_BEACON 0e570c3406b79266b7ada12e3b9314e7bb9dde3e    d78434e22fe8f37c65b6a2521362e1ae28345417    ceda879594cdb0c4874d05f68ae7f803404591e6    -75 -96 6.145786958466977   210bc1aa9f0fe9e2a60472f6d5ec9a34543997071559699291287   TYPE_BEACON b1607a1d8a0371d61472032c7562886c35337be6    7124f2ddfb9409f7830bad03ffb7347fa7d0ed08    0884d380b78ba64572be09f672702ddaed5b935f    -65 -89 10.256966296653632  b04aeb8aa364f0cda557e9365622dd7e11a98ddf1559699291695   TYPE_ACCELEROMETER  0.49412537  3.5641174   10.169495<br>\n\". </p>\n<p>SO we can see that each of the 3 entries in this line have different timestamps. Thanks!</p>",
      "rawMarkdown": "Thanks for sharing! I did not get how do you deal with the timestamps. They are normally added before the wort \"TYPE\". As for example here: \n\n\"1559699290975\tTYPE_BEACON\t0e570c3406b79266b7ada12e3b9314e7bb9dde3e\td78434e22fe8f37c65b6a2521362e1ae28345417\tceda879594cdb0c4874d05f68ae7f803404591e6\t-75\t-96\t6.145786958466977\t210bc1aa9f0fe9e2a60472f6d5ec9a34543997071559699291287\tTYPE_BEACON\tb1607a1d8a0371d61472032c7562886c35337be6\t7124f2ddfb9409f7830bad03ffb7347fa7d0ed08\t0884d380b78ba64572be09f672702ddaed5b935f\t-65\t-89\t10.256966296653632\tb04aeb8aa364f0cda557e9365622dd7e11a98ddf1559699291695\tTYPE_ACCELEROMETER\t0.49412537\t3.5641174\t10.169495\n\". \n\nSO we can see that each of the 3 entries in this line have different timestamps. Thanks!",
      "replies": [
        {
          "id": 1256924,
          "postDate": "2021-03-30T11:22:05.320Z",
          "content": "<p>Hi<br>\nI don't deal with timestamps at all in this patch. Before this patch the whole line you provided would be fed to further processing (so we loose 2 records). After my path, there are 3 lines.<br>\nDisclaimer: I added \\t before TYPE in your example, I guess tab was missed in copying, I think in the initial files there are tabs before TYPE.</p>\n<pre><code>new_list = []\nline=\"1559699290975\\tTYPE_BEACON 0e570c3406b79266b7ada12e3b9314e7bb9dde3e d78434e22fe8f37c65b6a2521362e1ae28345417 ceda879594cdb0c4874d05f68ae7f803404591e6 -75 -96 6.145786958466977 210bc1aa9f0fe9e2a60472f6d5ec9a34543997071559699291287\\tTYPE_BEACON b1607a1d8a0371d61472032c7562886c35337be6 7124f2ddfb9409f7830bad03ffb7347fa7d0ed08 0884d380b78ba64572be09f672702ddaed5b935f -65 -89 10.256966296653632 b04aeb8aa364f0cda557e9365622dd7e11a98ddf1559699291695\\tTYPE_ACCELEROMETER 0.49412537 3.5641174 10.169495\"\nif (\"TYPE_\" in line[30:]):                 \n    a = [m.start(0) for m in re.finditer(r'\\d{13}\\tTYPE', line[30:])]\n    print(a)\n    new_list.append(line[0:30 + a[0]] )              \n    for i in range(len(a)-1):\n        new_list.append(line[30:][a[i]:a[i+1]])   \n\n    new_list.append(line[30:][a[-1]:])    \nelse:\n    new_list.append(line)    \nfor elem in  new_list:\n    print(elem)\n\n\nOutput:\n[185, 401]\n1559699290975    TYPE_BEACON 0e570c3406b79266b7ada12e3b9314e7bb9dde3e d78434e22fe8f37c65b6a2521362e1ae28345417 ceda879594cdb0c4874d05f68ae7f803404591e6 -75 -96 6.145786958466977 210bc1aa9f0fe9e2a60472f6d5ec9a3454399707\n1559699291287    TYPE_BEACON b1607a1d8a0371d61472032c7562886c35337be6 7124f2ddfb9409f7830bad03ffb7347fa7d0ed08 0884d380b78ba64572be09f672702ddaed5b935f -65 -89 10.256966296653632 b04aeb8aa364f0cda557e9365622dd7e11a98ddf\n1559699291695    TYPE_ACCELEROMETER 0.49412537 3.5641174 10.169495\n</code></pre>\n<p>Thanks</p>",
          "rawMarkdown": "Hi\nI don't deal with timestamps at all in this patch. Before this patch the whole line you provided would be fed to further processing (so we loose 2 records). After my path, there are 3 lines.\nDisclaimer: I added \\t before TYPE in your example, I guess tab was missed in copying, I think in the initial files there are tabs before TYPE.\n\n```\nnew_list = []\nline=\"1559699290975\\tTYPE_BEACON 0e570c3406b79266b7ada12e3b9314e7bb9dde3e d78434e22fe8f37c65b6a2521362e1ae28345417 ceda879594cdb0c4874d05f68ae7f803404591e6 -75 -96 6.145786958466977 210bc1aa9f0fe9e2a60472f6d5ec9a34543997071559699291287\\tTYPE_BEACON b1607a1d8a0371d61472032c7562886c35337be6 7124f2ddfb9409f7830bad03ffb7347fa7d0ed08 0884d380b78ba64572be09f672702ddaed5b935f -65 -89 10.256966296653632 b04aeb8aa364f0cda557e9365622dd7e11a98ddf1559699291695\\tTYPE_ACCELEROMETER 0.49412537 3.5641174 10.169495\"\nif (\"TYPE_\" in line[30:]):                 \n    a = [m.start(0) for m in re.finditer(r'\\d{13}\\tTYPE', line[30:])]\n    print(a)\n    new_list.append(line[0:30 + a[0]] )              \n    for i in range(len(a)-1):\n        new_list.append(line[30:][a[i]:a[i+1]])   \n\n    new_list.append(line[30:][a[-1]:])    \nelse:\n    new_list.append(line)    \nfor elem in  new_list:\n    print(elem)\n\n\nOutput:\n[185, 401]\n1559699290975\tTYPE_BEACON 0e570c3406b79266b7ada12e3b9314e7bb9dde3e d78434e22fe8f37c65b6a2521362e1ae28345417 ceda879594cdb0c4874d05f68ae7f803404591e6 -75 -96 6.145786958466977 210bc1aa9f0fe9e2a60472f6d5ec9a3454399707\n1559699291287\tTYPE_BEACON b1607a1d8a0371d61472032c7562886c35337be6 7124f2ddfb9409f7830bad03ffb7347fa7d0ed08 0884d380b78ba64572be09f672702ddaed5b935f -65 -89 10.256966296653632 b04aeb8aa364f0cda557e9365622dd7e11a98ddf\n1559699291695\tTYPE_ACCELEROMETER 0.49412537 3.5641174 10.169495\n```\n\nThanks"
        }
      ]
    },
    {
      "id": 1251402,
      "postDate": "2021-03-24T18:33:51.603Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 1251412,
          "postDate": "2021-03-24T18:40:12.897Z",
          "content": "<p>30 is something to miss the first pattern.</p>\n<p>Thank you</p>",
          "rawMarkdown": "30 is something to miss the first pattern.\n\nThank you"
        },
        {
          "id": 1251443,
          "postDate": "2021-03-24T19:13:55.773Z",
          "content": "<p>Sorry for deleting the comment because I thought that might be a silly question.<br>\nI refresh the page, then I see your reply.</p>\n<p>So 30 is just something to skip the first \"timestamp TYPE_\" pattern. Thank you.</p>",
          "rawMarkdown": "Sorry for deleting the comment because I thought that might be a silly question.\nI refresh the page, then I see your reply.\n\nSo 30 is just something to skip the first \"timestamp TYPE_\" pattern. Thank you.",
          "votes": 1
        },
        {
          "id": 1251452,
          "postDate": "2021-03-24T19:22:27.190Z",
          "content": "<p>:) yes, first pattern is in each line, so we need to skip it to search for 'bad' lines.</p>\n<p>p.s. your question is completely ok. you never know whether this is your misunderstanding or a bug in the code. so feel free to ask</p>\n<p>Good luck</p>",
          "rawMarkdown": ":) yes, first pattern is in each line, so we need to skip it to search for 'bad' lines.\n\np.s. your question is completely ok. you never know whether this is your misunderstanding or a bug in the code. so feel free to ask\n\nGood luck"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1256663,
      "author_name": "Fedor F.",
      "author_url": "",
      "post_date": "2021-03-30T06:21:14.207000",
      "content": "<p>Thanks for sharing! I did not get how do you deal with the timestamps. They are normally added before the wort \"TYPE\". As for example here: </p>\n<p>\"1559699290975    TYPE_BEACON 0e570c3406b79266b7ada12e3b9314e7bb9dde3e    d78434e22fe8f37c65b6a2521362e1ae28345417    ceda879594cdb0c4874d05f68ae7f803404591e6    -75 -96 6.145786958466977   210bc1aa9f0fe9e2a60472f6d5ec9a34543997071559699291287   TYPE_BEACON b1607a1d8a0371d61472032c7562886c35337be6    7124f2ddfb9409f7830bad03ffb7347fa7d0ed08    0884d380b78ba64572be09f672702ddaed5b935f    -65 -89 10.256966296653632  b04aeb8aa364f0cda557e9365622dd7e11a98ddf1559699291695   TYPE_ACCELEROMETER  0.49412537  3.5641174   10.169495<br>\n\". </p>\n<p>SO we can see that each of the 3 entries in this line have different timestamps. Thanks!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1256924,
          "author_name": "Evgeniya",
          "author_url": "",
          "post_date": "2021-03-30T11:22:05.320000",
          "content": "<p>Hi<br>\nI don't deal with timestamps at all in this patch. Before this patch the whole line you provided would be fed to further processing (so we loose 2 records). After my path, there are 3 lines.<br>\nDisclaimer: I added \\t before TYPE in your example, I guess tab was missed in copying, I think in the initial files there are tabs before TYPE.</p>\n<pre><code>new_list = []\nline=\"1559699290975\\tTYPE_BEACON 0e570c3406b79266b7ada12e3b9314e7bb9dde3e d78434e22fe8f37c65b6a2521362e1ae28345417 ceda879594cdb0c4874d05f68ae7f803404591e6 -75 -96 6.145786958466977 210bc1aa9f0fe9e2a60472f6d5ec9a34543997071559699291287\\tTYPE_BEACON b1607a1d8a0371d61472032c7562886c35337be6 7124f2ddfb9409f7830bad03ffb7347fa7d0ed08 0884d380b78ba64572be09f672702ddaed5b935f -65 -89 10.256966296653632 b04aeb8aa364f0cda557e9365622dd7e11a98ddf1559699291695\\tTYPE_ACCELEROMETER 0.49412537 3.5641174 10.169495\"\nif (\"TYPE_\" in line[30:]):                 \n    a = [m.start(0) for m in re.finditer(r'\\d{13}\\tTYPE', line[30:])]\n    print(a)\n    new_list.append(line[0:30 + a[0]] )              \n    for i in range(len(a)-1):\n        new_list.append(line[30:][a[i]:a[i+1]])   \n\n    new_list.append(line[30:][a[-1]:])    \nelse:\n    new_list.append(line)    \nfor elem in  new_list:\n    print(elem)\n\n\nOutput:\n[185, 401]\n1559699290975    TYPE_BEACON 0e570c3406b79266b7ada12e3b9314e7bb9dde3e d78434e22fe8f37c65b6a2521362e1ae28345417 ceda879594cdb0c4874d05f68ae7f803404591e6 -75 -96 6.145786958466977 210bc1aa9f0fe9e2a60472f6d5ec9a3454399707\n1559699291287    TYPE_BEACON b1607a1d8a0371d61472032c7562886c35337be6 7124f2ddfb9409f7830bad03ffb7347fa7d0ed08 0884d380b78ba64572be09f672702ddaed5b935f -65 -89 10.256966296653632 b04aeb8aa364f0cda557e9365622dd7e11a98ddf\n1559699291695    TYPE_ACCELEROMETER 0.49412537 3.5641174 10.169495\n</code></pre>\n<p>Thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1251402,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-24T18:33:51.603000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1251412,
          "author_name": "Evgeniya",
          "author_url": "",
          "post_date": "2021-03-24T18:40:12.897000",
          "content": "<p>30 is something to miss the first pattern.</p>\n<p>Thank you</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1251443,
          "author_name": "kirksu",
          "author_url": "",
          "post_date": "2021-03-24T19:13:55.773000",
          "content": "<p>Sorry for deleting the comment because I thought that might be a silly question.<br>\nI refresh the page, then I see your reply.</p>\n<p>So 30 is just something to skip the first \"timestamp TYPE_\" pattern. Thank you.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1251452,
          "author_name": "Evgeniya",
          "author_url": "",
          "post_date": "2021-03-24T19:22:27.190000",
          "content": "<p>:) yes, first pattern is in each line, so we need to skip it to search for 'bad' lines.</p>\n<p>p.s. your question is completely ok. you never know whether this is your misunderstanding or a bug in the code. so feel free to ask</p>\n<p>Good luck</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1235135": "Hi, \nas you know sometimes \\n are missed in data files. To use io_f.read_data_file (from the repository provided) in your code, the following patch could be useful, the whole file is available at [https://www.kaggle.com/egm108/fix-for-io-f-py](url)\n\n```\n import numpy as np\n-\n+import re\n\n @dataclass\n class ReadData:\n@@ -31,8 +31,20 @@ def read_data_file(data_filename):\n\n     with open(data_filename, 'r', encoding='utf-8') as file:\n         lines = file.readlines()\n-\n-    for line_data in lines:\n+\n+    new_list = []\n+    for line in lines:\n+        if (\"TYPE_\" in line[30:]):\n+            a = [m.start(0) for m in re.finditer(r'\\d{13}\\tTYPE', line[30:])]\n+            new_list.append(line[0:30 + a[0]] )\n+            for i in range(len(a)-1):\n+                new_list.append(line[30:][a[i]:a[i+1]])\n+\n+            new_list.append(line[30:][a[-1]:])\n+        else:\n+            new_list.append(line)\n+\n+    for line_data in new_list:\n         line_data = line_data.strip()\n         if not line_data or line_data[0] == '#':\n:\n\n```\n\nThanks",
    "1256663": "Thanks for sharing! I did not get how do you deal with the timestamps. They are normally added before the wort \"TYPE\". As for example here: \n\n\"1559699290975\tTYPE_BEACON\t0e570c3406b79266b7ada12e3b9314e7bb9dde3e\td78434e22fe8f37c65b6a2521362e1ae28345417\tceda879594cdb0c4874d05f68ae7f803404591e6\t-75\t-96\t6.145786958466977\t210bc1aa9f0fe9e2a60472f6d5ec9a34543997071559699291287\tTYPE_BEACON\tb1607a1d8a0371d61472032c7562886c35337be6\t7124f2ddfb9409f7830bad03ffb7347fa7d0ed08\t0884d380b78ba64572be09f672702ddaed5b935f\t-65\t-89\t10.256966296653632\tb04aeb8aa364f0cda557e9365622dd7e11a98ddf1559699291695\tTYPE_ACCELEROMETER\t0.49412537\t3.5641174\t10.169495\n\". \n\nSO we can see that each of the 3 entries in this line have different timestamps. Thanks!",
    "1251402": ""
  }
}