{
  "id": 215604,
  "title": "How to solve data quality problem",
  "url": "/competitions/indoor-location-navigation/discussion/215604",
  "author_name": "Quvotha",
  "post_date": "2021-01-30T14:36:35.482000",
  "votes": 14,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I published following dataset. <br>\n<a href=\"https://www.kaggle.com/tomokikmogura/indoor-location-navigation-path-files\" target=\"_blank\">https://www.kaggle.com/tomokikmogura/indoor-location-navigation-path-files</a></p>\n<p>This dataset contains competition's path files preprocessed. I tried to make path files more handy and address data quality problem.  </p>\n<ul>\n<li>Format improved: separate each .txt path file into 3 files, header, footer, and body.<ul>\n<li>Body: .parquet file which is lighter than csv and can be read easily by <code>pandas.read_parquet()</code>.</li>\n<li>Footor and header: .json file which can be read easily by <code>json.load()</code>.</li></ul></li>\n<li>Data quality problem is resolved. </li>\n</ul>\n<p>The data quality problem is described as follows <a href=\"https://www.kaggle.com/c/indoor-location-navigation/data\" target=\"_blank\">here</a>.</p>\n<blockquote>\n  <p>A note on data quality: In the training files, you may find occasionally that a line is missing the ending newline character, causing it to run on to the next line. It is up to you how you want to handle this issue. This issue is not found in the test data.  </p>\n</blockquote>\n<p>This dataset is created by <a href=\"https://www.kaggle.com/tomokikmogura/convert-path-files-with-addressing-quality-problem\" target=\"_blank\">this notebook</a>.</p>\n<p>If you find something, please post a comment!  </p>\n<p>Thank you for reading.</p>",
  "messages": [
    {
      "id": 1177828,
      "postDate": "2021-01-30T14:36:35.483Z",
      "content": "<p>I published following dataset. <br>\n<a href=\"https://www.kaggle.com/tomokikmogura/indoor-location-navigation-path-files\" target=\"_blank\">https://www.kaggle.com/tomokikmogura/indoor-location-navigation-path-files</a></p>\n<p>This dataset contains competition's path files preprocessed. I tried to make path files more handy and address data quality problem.  </p>\n<ul>\n<li>Format improved: separate each .txt path file into 3 files, header, footer, and body.<ul>\n<li>Body: .parquet file which is lighter than csv and can be read easily by <code>pandas.read_parquet()</code>.</li>\n<li>Footor and header: .json file which can be read easily by <code>json.load()</code>.</li></ul></li>\n<li>Data quality problem is resolved. </li>\n</ul>\n<p>The data quality problem is described as follows <a href=\"https://www.kaggle.com/c/indoor-location-navigation/data\" target=\"_blank\">here</a>.</p>\n<blockquote>\n  <p>A note on data quality: In the training files, you may find occasionally that a line is missing the ending newline character, causing it to run on to the next line. It is up to you how you want to handle this issue. This issue is not found in the test data.  </p>\n</blockquote>\n<p>This dataset is created by <a href=\"https://www.kaggle.com/tomokikmogura/convert-path-files-with-addressing-quality-problem\" target=\"_blank\">this notebook</a>.</p>\n<p>If you find something, please post a comment!  </p>\n<p>Thank you for reading.</p>",
      "rawMarkdown": "I published following dataset. \nhttps://www.kaggle.com/tomokikmogura/indoor-location-navigation-path-files\n\nThis dataset contains competition's path files preprocessed. I tried to make path files more handy and address data quality problem.  \n\n- Format improved: separate each .txt path file into 3 files, header, footer, and body.\n  - Body: .parquet file which is lighter than csv and can be read easily by `pandas.read_parquet()`.\n  - Footor and header: .json file which can be read easily by `json.load()`.\n- Data quality problem is resolved. \n\nThe data quality problem is described as follows [here](https://www.kaggle.com/c/indoor-location-navigation/data).\n\n>A note on data quality: In the training files, you may find occasionally that a line is missing the ending newline character, causing it to run on to the next line. It is up to you how you want to handle this issue. This issue is not found in the test data.  \n\nThis dataset is created by [this notebook](https://www.kaggle.com/tomokikmogura/convert-path-files-with-addressing-quality-problem).\n\nIf you find something, please post a comment!  \n\nThank you for reading.",
      "votes": 13
    },
    {
      "id": 1240524,
      "postDate": "2021-03-16T13:11:12.607Z",
      "content": "<p>Thanks for sharing! Have you counted the number of files in which there is at least one '\\n' missing?</p>",
      "rawMarkdown": "Thanks for sharing! Have you counted the number of files in which there is at least one '\\n' missing?"
    },
    {
      "id": 1244137,
      "postDate": "2021-03-18T17:59:24.933Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1240524,
      "author_name": "Olaf Placha",
      "author_url": "",
      "post_date": "2021-03-16T13:11:12.607000",
      "content": "<p>Thanks for sharing! Have you counted the number of files in which there is at least one '\\n' missing?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1244137,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-18T17:59:24.933000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1177828": "I published following dataset. \nhttps://www.kaggle.com/tomokikmogura/indoor-location-navigation-path-files\n\nThis dataset contains competition's path files preprocessed. I tried to make path files more handy and address data quality problem.  \n\n- Format improved: separate each .txt path file into 3 files, header, footer, and body.\n  - Body: .parquet file which is lighter than csv and can be read easily by `pandas.read_parquet()`.\n  - Footor and header: .json file which can be read easily by `json.load()`.\n- Data quality problem is resolved. \n\nThe data quality problem is described as follows [here](https://www.kaggle.com/c/indoor-location-navigation/data).\n\n>A note on data quality: In the training files, you may find occasionally that a line is missing the ending newline character, causing it to run on to the next line. It is up to you how you want to handle this issue. This issue is not found in the test data.  \n\nThis dataset is created by [this notebook](https://www.kaggle.com/tomokikmogura/convert-path-files-with-addressing-quality-problem).\n\nIf you find something, please post a comment!  \n\nThank you for reading.",
    "1240524": "Thanks for sharing! Have you counted the number of files in which there is at least one '\\n' missing?",
    "1244137": ""
  }
}