{
  "id": 372990,
  "title": "How did you read the data?",
  "url": "/competitions/otto-recommender-system/discussion/372990",
  "author_name": "",
  "post_date": "2022-12-19T04:26:00.664364900Z",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I have recently started studying machine learning and would like to load the train data for this competition.<br>\nHowever, my computer has low specs, so if I use VS Code, the application will terminate halfway through.<br>\nAlso, if I use Google Colab, I run out of system RAM and lose the session.<br>\nHow did you guys load the train data?<br>\nI would appreciate it if you could let me know.</p>",
  "messages": [
    {
      "id": "2069520",
      "postDate": "12/19/2022 04:26:00",
      "content": "<p>I have recently started studying machine learning and would like to load the train data for this competition.<br>\nHowever, my computer has low specs, so if I use VS Code, the application will terminate halfway through.<br>\nAlso, if I use Google Colab, I run out of system RAM and lose the session.<br>\nHow did you guys load the train data?<br>\nI would appreciate it if you could let me know.</p>",
      "rawMarkdown": "I have recently started studying machine learning and would like to load the train data for this competition.\nHowever, my computer has low specs, so if I use VS Code, the application will terminate halfway through.\nAlso, if I use Google Colab, I run out of system RAM and lose the session.\nHow did you guys load the train data?\nI would appreciate it if you could let me know.",
      "votes": null
    },
    {
      "id": "2070311",
      "postDate": "12/19/2022 20:26:43",
      "content": "<p>If you use a Kaggle kernel, you should at least be able to read in the data (also consider using <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>'s processed parquet data <a href=\"https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint\" target=\"_blank\">here</a> which has a lower memory footprint).</p>",
      "rawMarkdown": "If you use a Kaggle kernel, you should at least be able to read in the data (also consider using @radek1's processed parquet data [here](https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint) which has a lower memory footprint).",
      "votes": null
    },
    {
      "id": "2070684",
      "postDate": "12/20/2022 09:06:09",
      "content": "<p>I decided to use Kaggle kernel!<br>\nAlso, the parquet file format is a new discovery for me!<br>\nAbhishek Shah, thank you for answering my question is my poor English!</p>",
      "rawMarkdown": "I decided to use Kaggle kernel!\nAlso, the parquet file format is a new discovery for me!\nAbhishek Shah, thank you for answering my question is my poor English!",
      "votes": null
    },
    {
      "id": "2070923",
      "postDate": "12/20/2022 13:21:49",
      "content": "<p>As an alternative approach I have created the following PySpark script, which reads-in the JSON files, converts them to a tabular PySpark dataframe and outputs the files as parquet. Using Spark, this code is able to run with a simple Kaggle kernel (in about an hour) w/o running into memory issues.<br>\n<a href=\"https://www.kaggle.com/code/ikogias/pre-process-with-pyspark-into-parquet\" target=\"_blank\">https://www.kaggle.com/code/ikogias/pre-process-with-pyspark-into-parquet</a></p>",
      "rawMarkdown": "As an alternative approach I have created the following PySpark script, which reads-in the JSON files, converts them to a tabular PySpark dataframe and outputs the files as parquet. Using Spark, this code is able to run with a simple Kaggle kernel (in about an hour) w/o running into memory issues.\n\nhttps://www.kaggle.com/code/ikogias/pre-process-with-pyspark-into-parquet",
      "votes": null
    },
    {
      "id": "2072394",
      "postDate": "12/22/2022 04:36:18",
      "content": "<p>Thank you, <a href=\"https://www.kaggle.com/ikogias\" target=\"_blank\">@ikogias</a> !<br>\nPySpark is a new discovery for me!<br>\nI'm going to look into PySpark!</p>",
      "rawMarkdown": "Thank you, @ikogias !\nPySpark is a new discovery for me!\nI'm going to look into PySpark!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2070311,
      "author_name": "alberteinsten",
      "author_url": "",
      "post_date": "12/19/2022 20:26:43",
      "content": "<p>If you use a Kaggle kernel, you should at least be able to read in the data (also consider using <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>'s processed parquet data <a href=\"https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint\" target=\"_blank\">here</a> which has a lower memory footprint).</p>",
      "votes": null,
      "replies": [
        {
          "id": 2070684,
          "author_name": "takuma0306",
          "author_url": "",
          "post_date": "12/20/2022 09:06:09",
          "content": "<p>I decided to use Kaggle kernel!<br>\nAlso, the parquet file format is a new discovery for me!<br>\nAbhishek Shah, thank you for answering my question is my poor English!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2070923,
      "author_name": "ikogias",
      "author_url": "",
      "post_date": "12/20/2022 13:21:49",
      "content": "<p>As an alternative approach I have created the following PySpark script, which reads-in the JSON files, converts them to a tabular PySpark dataframe and outputs the files as parquet. Using Spark, this code is able to run with a simple Kaggle kernel (in about an hour) w/o running into memory issues.<br>\n<a href=\"https://www.kaggle.com/code/ikogias/pre-process-with-pyspark-into-parquet\" target=\"_blank\">https://www.kaggle.com/code/ikogias/pre-process-with-pyspark-into-parquet</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2072394,
          "author_name": "takuma0306",
          "author_url": "",
          "post_date": "12/22/2022 04:36:18",
          "content": "<p>Thank you, <a href=\"https://www.kaggle.com/ikogias\" target=\"_blank\">@ikogias</a> !<br>\nPySpark is a new discovery for me!<br>\nI'm going to look into PySpark!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2069520": "I have recently started studying machine learning and would like to load the train data for this competition.\nHowever, my computer has low specs, so if I use VS Code, the application will terminate halfway through.\nAlso, if I use Google Colab, I run out of system RAM and lose the session.\nHow did you guys load the train data?\nI would appreciate it if you could let me know.",
    "2070311": "If you use a Kaggle kernel, you should at least be able to read in the data (also consider using @radek1's processed parquet data [here](https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint) which has a lower memory footprint).",
    "2070684": "I decided to use Kaggle kernel!\nAlso, the parquet file format is a new discovery for me!\nAbhishek Shah, thank you for answering my question is my poor English!",
    "2070923": "As an alternative approach I have created the following PySpark script, which reads-in the JSON files, converts them to a tabular PySpark dataframe and outputs the files as parquet. Using Spark, this code is able to run with a simple Kaggle kernel (in about an hour) w/o running into memory issues.\n\nhttps://www.kaggle.com/code/ikogias/pre-process-with-pyspark-into-parquet",
    "2072394": "Thank you, @ikogias !\nPySpark is a new discovery for me!\nI'm going to look into PySpark!"
  },
  "source": "meta"
}