{
  "id": 228327,
  "title": "How to fix malformed lines in train & test data",
  "url": "/competitions/indoor-location-navigation/discussion/228327",
  "author_name": "",
  "post_date": "2021-03-24T07:21:01.510227900Z",
  "votes": 16,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Some people have reported there are malformed lines in train &amp; test data. eg) <a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/215973\" target=\"_blank\">Some of the txt data are broken</a> by <a href=\"https://www.kaggle.com/kenmatsu4\" target=\"_blank\">@kenmatsu4</a>. There are some lines with multiple TYPE_ in the files. </p>\n<p>Example)<br>\n<img src=\"https://user-images.githubusercontent.com/54491/112269107-45a2c400-8cbb-11eb-9d78-59b0254643a6.png\" alt=\"malformed\"></p>\n<p>There are already great workaround shared as follows.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/kenmatsu4/feature-store-for-indoor-location-navigation\" target=\"_blank\">feature store for Indoor Location &amp; Navigation</a></li>\n<li><a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/225353\" target=\"_blank\">fix for host's io_f.py to handle missed new lines</a></li>\n</ul>\n<p>But I'm introducing new one just because I run my data processing a lot. If I fix the lines when <em>read time</em>, data processing becomes a lot slower. To mitigate that I decided to fix the files only once in place.</p>\n<h3>How we do it.</h3>\n<p>In short you just need to run<br>\n<code>find ./ -type f -exec sed -i -e 's/[^^]\\([[:digit:]]\\{13\\}\\)\\tTYPE/\\n\\1\\tTYPE/g' {} \\;</code><br>\nin your data directory. This will fix the files in place. Note that it will take a few hours.</p>\n<p>Please see <a href=\"https://www.kaggle.com/higepon/how-to-fix-malformed-train-test-data\" target=\"_blank\">my notebook</a> for more details.</p>\n<h3>Disclaimer</h3>\n<p>I maybe making silly mistake in the command. Please do it at your own risk :)<br>\nAny feedback appreciated.</p>\n<p>I wish I could share the fixed dataset but the dataset is too big to create kaggle dataset unfortunately.</p>",
  "messages": [
    {
      "id": "1250644",
      "postDate": "03/24/2021 07:21:01",
      "content": "<p>Some people have reported there are malformed lines in train &amp; test data. eg) <a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/215973\" target=\"_blank\">Some of the txt data are broken</a> by <a href=\"https://www.kaggle.com/kenmatsu4\" target=\"_blank\">@kenmatsu4</a>. There are some lines with multiple TYPE_ in the files. </p>\n<p>Example)<br>\n<img src=\"https://user-images.githubusercontent.com/54491/112269107-45a2c400-8cbb-11eb-9d78-59b0254643a6.png\" alt=\"malformed\"></p>\n<p>There are already great workaround shared as follows.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/kenmatsu4/feature-store-for-indoor-location-navigation\" target=\"_blank\">feature store for Indoor Location &amp; Navigation</a></li>\n<li><a href=\"https://www.kaggle.com/c/indoor-location-navigation/discussion/225353\" target=\"_blank\">fix for host's io_f.py to handle missed new lines</a></li>\n</ul>\n<p>But I'm introducing new one just because I run my data processing a lot. If I fix the lines when <em>read time</em>, data processing becomes a lot slower. To mitigate that I decided to fix the files only once in place.</p>\n<h3>How we do it.</h3>\n<p>In short you just need to run<br>\n<code>find ./ -type f -exec sed -i -e 's/[^^]\\([[:digit:]]\\{13\\}\\)\\tTYPE/\\n\\1\\tTYPE/g' {} \\;</code><br>\nin your data directory. This will fix the files in place. Note that it will take a few hours.</p>\n<p>Please see <a href=\"https://www.kaggle.com/higepon/how-to-fix-malformed-train-test-data\" target=\"_blank\">my notebook</a> for more details.</p>\n<h3>Disclaimer</h3>\n<p>I maybe making silly mistake in the command. Please do it at your own risk :)<br>\nAny feedback appreciated.</p>\n<p>I wish I could share the fixed dataset but the dataset is too big to create kaggle dataset unfortunately.</p>",
      "rawMarkdown": "Some people have reported there are malformed lines in train & test data. eg) [Some of the txt data are broken](https://www.kaggle.com/c/indoor-location-navigation/discussion/215973) by @kenmatsu4. There are some lines with multiple TYPE_ in the files. \n\nExample)\n![malformed](https://user-images.githubusercontent.com/54491/112269107-45a2c400-8cbb-11eb-9d78-59b0254643a6.png)\n\nThere are already great workaround shared as follows.\n- [feature store for Indoor Location & Navigation](https://www.kaggle.com/kenmatsu4/feature-store-for-indoor-location-navigation)\n- [fix for host's io_f.py to handle missed new lines](https://www.kaggle.com/c/indoor-location-navigation/discussion/225353)\n\nBut I'm introducing new one just because I run my data processing a lot. If I fix the lines when *read time*, data processing becomes a lot slower. To mitigate that I decided to fix the files only once in place.\n\n### How we do it.\nIn short you just need to run\n```find ./ -type f -exec sed -i -e 's/[^^]\\([[:digit:]]\\{13\\}\\)\\tTYPE/\\n\\1\\tTYPE/g' {} \\;```\nin your data directory. This will fix the files in place. Note that it will take a few hours.\n\nPlease see [my notebook](https://www.kaggle.com/higepon/how-to-fix-malformed-train-test-data) for more details.\n### Disclaimer\nI maybe making silly mistake in the command. Please do it at your own risk :)\nAny feedback appreciated.\n\n\nI wish I could share the fixed dataset but the dataset is too big to create kaggle dataset unfortunately.",
      "votes": null
    },
    {
      "id": "1277992",
      "postDate": "04/19/2021 13:13:19",
      "content": "<p>Thank you for the helpful script!</p>\n<p>Unfortunately, I found some minor issues, so I want to share my findings here.</p>\n<p><strong>1. <code>s/[^^].../.../g</code> should be <code>s/\\([^^]\\).../\\1.../g</code>.</strong></p>\n<p>Seeing the diff result carefully, I found the last hash field of resulting BEACON_TYPE records was missing the last character.<br>\nThis is an example of the diff.</p>\n<pre><code>-1559714239537    TYPE_BEACON .... 02e5d6144f3d7fe30c23af14973709ba2c1b1ce81559714240324  TYPE_ACCELEROMETER  -0.44984436 1.6701965   9.7005005\n+1559714239537    TYPE_BEACON .... 02e5d6144f3d7fe30c23af14973709ba2c1b1ce\n+1559714240324    TYPE_ACCELEROMETER  -0.44984436 1.6701965   9.7005005\n</code></pre>\n<p>In the above example, <code>02e5d6144f3d7fe30c23af14973709ba2c1b1ce81559714240324</code> was split into <br>\n<code>02e5d6144f3d7fe30c23af14973709ba2c1b1ce</code>  and <code>1559714240324</code>. But it should be<br>\n<code>02e5d6144f3d7fe30c23af14973709ba2c1b1ce8</code> and <code>1559714240324</code>.</p>\n<p>When I found it at first, I couldn't figure out why 🤔.<br>\nBut I found that the <code>[^^]</code> itself matched a character.</p>\n<p>This is a simple demo:</p>\n<pre><code>$ echo abababab | sed -e 's/[^^]a/A/g'\naAAAb\n</code></pre>\n<p>So, we need to capture the results of <code>[^^]</code> explicitly and restore them.</p>\n<pre><code>$ echo abababab | sed -e 's/\\([^^]\\)a/\\1A/g'\nabAbAbAb\n</code></pre>\n<p><strong>2. TYPE_BEACON data of the broken records are missing the last timestamp.</strong></p>\n<p>Note that this is not the issue of the script but the original data problem.</p>\n<p>I am not sure this can affect the model score, but making the data consistent might be desirable.<br>\nSince the TYPE_BEACON's last field is just a copy of the first field timestamp, it can be repaired by capturing the first field.</p>\n<pre><code>sed -e 's/\\([[:digit:]]\\{13\\}\\)\\(\\tTYPE_BEACON.*[0-9a-f]\\{40\\}\\)$/\\1\\2\\t\\1/g';\n</code></pre>\n<p><strong>Idea of Fix</strong></p>\n<p>I could get results including the above fixes by applying the following two commands.</p>\n<pre><code>find ./ -type f -exec sed -i -e 's/\\([^^]\\)\\([[:digit:]]\\{13\\}\\)\\tTYPE/\\1\\n\\2\\tTYPE/g' {} \\;\nfind ./ -type f -exec sed -i -e 's/\\([[:digit:]]\\{13\\}\\)\\(\\tTYPE_BEACON.*[0-9a-f]\\{40\\}\\)$/\\1\\2\\t\\1/g' {} \\;\n</code></pre>\n<p>Note that it will take two times longer than the original one. 😅<br>\nThe second command might be optional because we can use the first field timestamp instead of the last one for the beacon data without the extra computational cost.</p>",
      "rawMarkdown": "Thank you for the helpful script!\n\nUnfortunately, I found some minor issues, so I want to share my findings here.\n\n**1. `s/[^^].../.../g` should be `s/\\([^^]\\).../\\1.../g`.**\n\nSeeing the diff result carefully, I found the last hash field of resulting BEACON_TYPE records was missing the last character.\nThis is an example of the diff.\n\n```\n-1559714239537\tTYPE_BEACON\t.... 02e5d6144f3d7fe30c23af14973709ba2c1b1ce81559714240324\tTYPE_ACCELEROMETER\t-0.44984436\t1.6701965\t9.7005005\n+1559714239537\tTYPE_BEACON\t.... 02e5d6144f3d7fe30c23af14973709ba2c1b1ce\n+1559714240324\tTYPE_ACCELEROMETER\t-0.44984436\t1.6701965\t9.7005005\n```\n\nIn the above example, `02e5d6144f3d7fe30c23af14973709ba2c1b1ce81559714240324` was split into \n`02e5d6144f3d7fe30c23af14973709ba2c1b1ce`  and `1559714240324`. But it should be\n`02e5d6144f3d7fe30c23af14973709ba2c1b1ce8` and `1559714240324`.\n\nWhen I found it at first, I couldn't figure out why 🤔.\nBut I found that the `[^^]` itself matched a character.\n\nThis is a simple demo:\n\n```\n$ echo abababab | sed -e 's/[^^]a/A/g'\naAAAb\n```\n\nSo, we need to capture the results of `[^^]` explicitly and restore them.\n\n```\n$ echo abababab | sed -e 's/\\([^^]\\)a/\\1A/g'\nabAbAbAb\n```\n\n**2. TYPE_BEACON data of the broken records are missing the last timestamp.**\n\nNote that this is not the issue of the script but the original data problem.\n\nI am not sure this can affect the model score, but making the data consistent might be desirable.\nSince the TYPE_BEACON's last field is just a copy of the first field timestamp, it can be repaired by capturing the first field.\n\n```\nsed -e 's/\\([[:digit:]]\\{13\\}\\)\\(\\tTYPE_BEACON.*[0-9a-f]\\{40\\}\\)$/\\1\\2\\t\\1/g';\n```\n\n**Idea of Fix**\n\nI could get results including the above fixes by applying the following two commands.\n\n```\nfind ./ -type f -exec sed -i -e 's/\\([^^]\\)\\([[:digit:]]\\{13\\}\\)\\tTYPE/\\1\\n\\2\\tTYPE/g' {} \\;\nfind ./ -type f -exec sed -i -e 's/\\([[:digit:]]\\{13\\}\\)\\(\\tTYPE_BEACON.*[0-9a-f]\\{40\\}\\)$/\\1\\2\\t\\1/g' {} \\;\n```\n\nNote that it will take two times longer than the original one. 😅\nThe second command might be optional because we can use the first field timestamp instead of the last one for the beacon data without the extra computational cost.",
      "votes": null
    },
    {
      "id": "1278474",
      "postDate": "04/19/2021 23:53:40",
      "content": "<p>Thanks for catching the issue and sharing the fix. I'll apply your fix.</p>",
      "rawMarkdown": "Thanks for catching the issue and sharing the fix. I'll apply your fix.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1277992,
      "author_name": "ishicamo",
      "author_url": "",
      "post_date": "04/19/2021 13:13:19",
      "content": "<p>Thank you for the helpful script!</p>\n<p>Unfortunately, I found some minor issues, so I want to share my findings here.</p>\n<p><strong>1. <code>s/[^^].../.../g</code> should be <code>s/\\([^^]\\).../\\1.../g</code>.</strong></p>\n<p>Seeing the diff result carefully, I found the last hash field of resulting BEACON_TYPE records was missing the last character.<br>\nThis is an example of the diff.</p>\n<pre><code>-1559714239537    TYPE_BEACON .... 02e5d6144f3d7fe30c23af14973709ba2c1b1ce81559714240324  TYPE_ACCELEROMETER  -0.44984436 1.6701965   9.7005005\n+1559714239537    TYPE_BEACON .... 02e5d6144f3d7fe30c23af14973709ba2c1b1ce\n+1559714240324    TYPE_ACCELEROMETER  -0.44984436 1.6701965   9.7005005\n</code></pre>\n<p>In the above example, <code>02e5d6144f3d7fe30c23af14973709ba2c1b1ce81559714240324</code> was split into <br>\n<code>02e5d6144f3d7fe30c23af14973709ba2c1b1ce</code>  and <code>1559714240324</code>. But it should be<br>\n<code>02e5d6144f3d7fe30c23af14973709ba2c1b1ce8</code> and <code>1559714240324</code>.</p>\n<p>When I found it at first, I couldn't figure out why 🤔.<br>\nBut I found that the <code>[^^]</code> itself matched a character.</p>\n<p>This is a simple demo:</p>\n<pre><code>$ echo abababab | sed -e 's/[^^]a/A/g'\naAAAb\n</code></pre>\n<p>So, we need to capture the results of <code>[^^]</code> explicitly and restore them.</p>\n<pre><code>$ echo abababab | sed -e 's/\\([^^]\\)a/\\1A/g'\nabAbAbAb\n</code></pre>\n<p><strong>2. TYPE_BEACON data of the broken records are missing the last timestamp.</strong></p>\n<p>Note that this is not the issue of the script but the original data problem.</p>\n<p>I am not sure this can affect the model score, but making the data consistent might be desirable.<br>\nSince the TYPE_BEACON's last field is just a copy of the first field timestamp, it can be repaired by capturing the first field.</p>\n<pre><code>sed -e 's/\\([[:digit:]]\\{13\\}\\)\\(\\tTYPE_BEACON.*[0-9a-f]\\{40\\}\\)$/\\1\\2\\t\\1/g';\n</code></pre>\n<p><strong>Idea of Fix</strong></p>\n<p>I could get results including the above fixes by applying the following two commands.</p>\n<pre><code>find ./ -type f -exec sed -i -e 's/\\([^^]\\)\\([[:digit:]]\\{13\\}\\)\\tTYPE/\\1\\n\\2\\tTYPE/g' {} \\;\nfind ./ -type f -exec sed -i -e 's/\\([[:digit:]]\\{13\\}\\)\\(\\tTYPE_BEACON.*[0-9a-f]\\{40\\}\\)$/\\1\\2\\t\\1/g' {} \\;\n</code></pre>\n<p>Note that it will take two times longer than the original one. 😅<br>\nThe second command might be optional because we can use the first field timestamp instead of the last one for the beacon data without the extra computational cost.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1278474,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "04/19/2021 23:53:40",
          "content": "<p>Thanks for catching the issue and sharing the fix. I'll apply your fix.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1250644": "Some people have reported there are malformed lines in train & test data. eg) [Some of the txt data are broken](https://www.kaggle.com/c/indoor-location-navigation/discussion/215973) by @kenmatsu4. There are some lines with multiple TYPE_ in the files. \n\nExample)\n![malformed](https://user-images.githubusercontent.com/54491/112269107-45a2c400-8cbb-11eb-9d78-59b0254643a6.png)\n\nThere are already great workaround shared as follows.\n- [feature store for Indoor Location & Navigation](https://www.kaggle.com/kenmatsu4/feature-store-for-indoor-location-navigation)\n- [fix for host's io_f.py to handle missed new lines](https://www.kaggle.com/c/indoor-location-navigation/discussion/225353)\n\nBut I'm introducing new one just because I run my data processing a lot. If I fix the lines when *read time*, data processing becomes a lot slower. To mitigate that I decided to fix the files only once in place.\n\n### How we do it.\nIn short you just need to run\n```find ./ -type f -exec sed -i -e 's/[^^]\\([[:digit:]]\\{13\\}\\)\\tTYPE/\\n\\1\\tTYPE/g' {} \\;```\nin your data directory. This will fix the files in place. Note that it will take a few hours.\n\nPlease see [my notebook](https://www.kaggle.com/higepon/how-to-fix-malformed-train-test-data) for more details.\n### Disclaimer\nI maybe making silly mistake in the command. Please do it at your own risk :)\nAny feedback appreciated.\n\n\nI wish I could share the fixed dataset but the dataset is too big to create kaggle dataset unfortunately.",
    "1277992": "Thank you for the helpful script!\n\nUnfortunately, I found some minor issues, so I want to share my findings here.\n\n**1. `s/[^^].../.../g` should be `s/\\([^^]\\).../\\1.../g`.**\n\nSeeing the diff result carefully, I found the last hash field of resulting BEACON_TYPE records was missing the last character.\nThis is an example of the diff.\n\n```\n-1559714239537\tTYPE_BEACON\t.... 02e5d6144f3d7fe30c23af14973709ba2c1b1ce81559714240324\tTYPE_ACCELEROMETER\t-0.44984436\t1.6701965\t9.7005005\n+1559714239537\tTYPE_BEACON\t.... 02e5d6144f3d7fe30c23af14973709ba2c1b1ce\n+1559714240324\tTYPE_ACCELEROMETER\t-0.44984436\t1.6701965\t9.7005005\n```\n\nIn the above example, `02e5d6144f3d7fe30c23af14973709ba2c1b1ce81559714240324` was split into \n`02e5d6144f3d7fe30c23af14973709ba2c1b1ce`  and `1559714240324`. But it should be\n`02e5d6144f3d7fe30c23af14973709ba2c1b1ce8` and `1559714240324`.\n\nWhen I found it at first, I couldn't figure out why 🤔.\nBut I found that the `[^^]` itself matched a character.\n\nThis is a simple demo:\n\n```\n$ echo abababab | sed -e 's/[^^]a/A/g'\naAAAb\n```\n\nSo, we need to capture the results of `[^^]` explicitly and restore them.\n\n```\n$ echo abababab | sed -e 's/\\([^^]\\)a/\\1A/g'\nabAbAbAb\n```\n\n**2. TYPE_BEACON data of the broken records are missing the last timestamp.**\n\nNote that this is not the issue of the script but the original data problem.\n\nI am not sure this can affect the model score, but making the data consistent might be desirable.\nSince the TYPE_BEACON's last field is just a copy of the first field timestamp, it can be repaired by capturing the first field.\n\n```\nsed -e 's/\\([[:digit:]]\\{13\\}\\)\\(\\tTYPE_BEACON.*[0-9a-f]\\{40\\}\\)$/\\1\\2\\t\\1/g';\n```\n\n**Idea of Fix**\n\nI could get results including the above fixes by applying the following two commands.\n\n```\nfind ./ -type f -exec sed -i -e 's/\\([^^]\\)\\([[:digit:]]\\{13\\}\\)\\tTYPE/\\1\\n\\2\\tTYPE/g' {} \\;\nfind ./ -type f -exec sed -i -e 's/\\([[:digit:]]\\{13\\}\\)\\(\\tTYPE_BEACON.*[0-9a-f]\\{40\\}\\)$/\\1\\2\\t\\1/g' {} \\;\n```\n\nNote that it will take two times longer than the original one. 😅\nThe second command might be optional because we can use the first field timestamp instead of the last one for the beacon data without the extra computational cost.",
    "1278474": "Thanks for catching the issue and sharing the fix. I'll apply your fix."
  },
  "source": "meta"
}