{
  "id": 366709,
  "title": "Summary about Loading and Preprocessing Big Jsonl Data File",
  "url": "/competitions/otto-recommender-system/discussion/366709",
  "author_name": "",
  "post_date": "2022-11-17T09:52:56.914588Z",
  "votes": 12,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Below are the notebooks that help me a lot with loading and preprocessing data for this competition.<br>\nI divided it into two parts: Load and preprocessing(convert to dataframe, csv or parquet):</p>\n<p>Notebooks [1] to [3] are about loading the jsonl files <br>\nNotebooks [4] to [6] are about the processing <br>\n[6] provided a new dataset which is parquet/csv files<br>\n<strong>If you like the below notebooks, please help to upvote after reading then</strong></p>\n<ol>\n<li><p><a href=\"https://www.kaggle.com/code/xxxxyyyy80008/otto-multi-objective-recommender-system-eda/notebook\" target=\"_blank\">Reading jsonl data using HuggingFace Datasets</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/code/edwardcrookenden/otto-getting-started-eda-baseline\" target=\"_blank\">📊 OTTO - Getting Started (EDA + Baseline)🧑‍💻</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/code/xxxxyyyy80008/how-to-read-a-big-json-file-with-python\" target=\"_blank\">How to read a big JSON file with Python</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/code/columbia2131/otto-read-a-chunk-of-jsonl-to-manageable-df\" target=\"_blank\">OTTO - Read a chunk of jsonl to manageable DF</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/code/columbia2131/otto-fast-dataframe-loading-in-parquet-format\" target=\"_blank\">OTTO - Fast DataFrame Loading in Parquet format</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/code/radek1/howto-full-dataset-as-parquet-csv-files?scriptVersionId=109945227\" target=\"_blank\">💡 [Howto] Full dataset as parquet/csv files</a></p></li>\n</ol>",
  "messages": [
    {
      "id": "2033498",
      "postDate": "11/17/2022 09:52:56",
      "content": "<p>Below are the notebooks that help me a lot with loading and preprocessing data for this competition.<br>\nI divided it into two parts: Load and preprocessing(convert to dataframe, csv or parquet):</p>\n<p>Notebooks [1] to [3] are about loading the jsonl files <br>\nNotebooks [4] to [6] are about the processing <br>\n[6] provided a new dataset which is parquet/csv files<br>\n<strong>If you like the below notebooks, please help to upvote after reading then</strong></p>\n<ol>\n<li><p><a href=\"https://www.kaggle.com/code/xxxxyyyy80008/otto-multi-objective-recommender-system-eda/notebook\" target=\"_blank\">Reading jsonl data using HuggingFace Datasets</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/code/edwardcrookenden/otto-getting-started-eda-baseline\" target=\"_blank\">📊 OTTO - Getting Started (EDA + Baseline)🧑‍💻</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/code/xxxxyyyy80008/how-to-read-a-big-json-file-with-python\" target=\"_blank\">How to read a big JSON file with Python</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/code/columbia2131/otto-read-a-chunk-of-jsonl-to-manageable-df\" target=\"_blank\">OTTO - Read a chunk of jsonl to manageable DF</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/code/columbia2131/otto-fast-dataframe-loading-in-parquet-format\" target=\"_blank\">OTTO - Fast DataFrame Loading in Parquet format</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/code/radek1/howto-full-dataset-as-parquet-csv-files?scriptVersionId=109945227\" target=\"_blank\">💡 [Howto] Full dataset as parquet/csv files</a></p></li>\n</ol>",
      "rawMarkdown": "Below are the notebooks that help me a lot with loading and preprocessing data for this competition.\nI divided it into two parts: Load and preprocessing(convert to dataframe, csv or parquet):\n\nNotebooks [1] to [3] are about loading the jsonl files \nNotebooks [4] to [6] are about the processing \n[6] provided a new dataset which is parquet/csv files\n**If you like the below notebooks, please help to upvote after reading then**\n\n\n1. [Reading jsonl data using HuggingFace Datasets](https://www.kaggle.com/code/xxxxyyyy80008/otto-multi-objective-recommender-system-eda/notebook)\n\n2. [📊 OTTO - Getting Started (EDA + Baseline)🧑‍💻](https://www.kaggle.com/code/edwardcrookenden/otto-getting-started-eda-baseline)\n\n3. [How to read a big JSON file with Python](https://www.kaggle.com/code/xxxxyyyy80008/how-to-read-a-big-json-file-with-python)\n\n4. [OTTO - Read a chunk of jsonl to manageable DF](https://www.kaggle.com/code/columbia2131/otto-read-a-chunk-of-jsonl-to-manageable-df)\n\n5.  [OTTO - Fast DataFrame Loading in Parquet format](https://www.kaggle.com/code/columbia2131/otto-fast-dataframe-loading-in-parquet-format)\n\n6. [💡 [Howto] Full dataset as parquet/csv files](https://www.kaggle.com/code/radek1/howto-full-dataset-as-parquet-csv-files?scriptVersionId=109945227)",
      "votes": null
    },
    {
      "id": "2033508",
      "postDate": "11/17/2022 09:59:57",
      "content": "<p>Thanks a lot for the info</p>",
      "rawMarkdown": "Thanks a lot for the info",
      "votes": null
    },
    {
      "id": "2056595",
      "postDate": "12/06/2022 09:26:30",
      "content": "<p>Thank you for reviewing it</p>",
      "rawMarkdown": "Thank you for reviewing it",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2033508,
      "author_name": "mridulsyed",
      "author_url": "",
      "post_date": "11/17/2022 09:59:57",
      "content": "<p>Thanks a lot for the info</p>",
      "votes": null,
      "replies": [
        {
          "id": 2056595,
          "author_name": "leiwong",
          "author_url": "",
          "post_date": "12/06/2022 09:26:30",
          "content": "<p>Thank you for reviewing it</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2033498": "Below are the notebooks that help me a lot with loading and preprocessing data for this competition.\nI divided it into two parts: Load and preprocessing(convert to dataframe, csv or parquet):\n\nNotebooks [1] to [3] are about loading the jsonl files \nNotebooks [4] to [6] are about the processing \n[6] provided a new dataset which is parquet/csv files\n**If you like the below notebooks, please help to upvote after reading then**\n\n\n1. [Reading jsonl data using HuggingFace Datasets](https://www.kaggle.com/code/xxxxyyyy80008/otto-multi-objective-recommender-system-eda/notebook)\n\n2. [📊 OTTO - Getting Started (EDA + Baseline)🧑‍💻](https://www.kaggle.com/code/edwardcrookenden/otto-getting-started-eda-baseline)\n\n3. [How to read a big JSON file with Python](https://www.kaggle.com/code/xxxxyyyy80008/how-to-read-a-big-json-file-with-python)\n\n4. [OTTO - Read a chunk of jsonl to manageable DF](https://www.kaggle.com/code/columbia2131/otto-read-a-chunk-of-jsonl-to-manageable-df)\n\n5.  [OTTO - Fast DataFrame Loading in Parquet format](https://www.kaggle.com/code/columbia2131/otto-fast-dataframe-loading-in-parquet-format)\n\n6. [💡 [Howto] Full dataset as parquet/csv files](https://www.kaggle.com/code/radek1/howto-full-dataset-as-parquet-csv-files?scriptVersionId=109945227)",
    "2033508": "Thanks a lot for the info",
    "2056595": "Thank you for reviewing it"
  },
  "source": "meta"
}