{
  "id": 121009,
  "title": "How to handle PlayerTrackData dataset in R",
  "url": "/competitions/nfl-playing-surface-analytics/discussion/121009",
  "author_name": "Harish Nagpal",
  "post_date": "2019-12-10T12:46:06.151000",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>All, any idea how to handle this large data set of 76 million records in R. I am using R and R hangs if I try to fetch data from this data set. \nThanks in advance.\nHowever if I do analysis in SQLite Studio, then it works although some what slow.\nI have 4 GB RAM on my PC.</p>",
  "messages": [
    {
      "id": 691741,
      "postDate": "2019-12-10T13:02:43.073Z",
      "content": "<p>Use Kaggle kernel dude <a href=\"/harnagpal\">@harnagpal</a>\nIt provides 16 GB RAM :) \n<a href=\"https://www.kaggle.com/docs/kernels\">https://www.kaggle.com/docs/kernels</a></p>",
      "rawMarkdown": "Use Kaggle kernel dude @harnagpal\nIt provides 16 GB RAM :) \nhttps://www.kaggle.com/docs/kernels",
      "votes": 1,
      "replies": [
        {
          "id": 691788,
          "postDate": "2019-12-10T13:54:19.393Z",
          "content": "<p>I would suggest the same.\nIf you however insist on using your local machine then you could partially load the dataset through data.table library's function fread() playing with parameters skip and nrow.</p>",
          "rawMarkdown": "I would suggest the same.\nIf you however insist on using your local machine then you could partially load the dataset through data.table library's function fread() playing with parameters skip and nrow.",
          "votes": 2
        },
        {
          "id": 692727,
          "postDate": "2019-12-11T17:05:21.427Z",
          "content": "<p>Thanks Rasyid, Kröger\nThat's great. I didn't know that we can use Kaggle like this.\nMeanwhile I tried with RSQLite pacakge. And it seems its working for me. I created a database and imported csv files into it in tables and then called them through data.table.</p>",
          "rawMarkdown": "Thanks Rasyid, Kröger\nThat's great. I didn't know that we can use Kaggle like this.\nMeanwhile I tried with RSQLite pacakge. And it seems its working for me. I created a database and imported csv files into it in tables and then called them through data.table."
        }
      ]
    },
    {
      "id": 691726,
      "postDate": "2019-12-10T12:46:06.153Z",
      "content": "<p>All, any idea how to handle this large data set of 76 million records in R. I am using R and R hangs if I try to fetch data from this data set. \nThanks in advance.\nHowever if I do analysis in SQLite Studio, then it works although some what slow.\nI have 4 GB RAM on my PC.</p>",
      "rawMarkdown": " All, any idea how to handle this large data set of 76 million records in R. I am using R and R hangs if I try to fetch data from this data set. \nThanks in advance.\nHowever if I do analysis in SQLite Studio, then it works although some what slow.\nI have 4 GB RAM on my PC.",
      "votes": 1
    },
    {
      "id": 694618,
      "postDate": "2019-12-13T21:10:59.677Z",
      "content": "<p>You cannot use R to handle that size of data, even loaded to Kaggle kernel.  And I find all of the public notebooks right now are simply playing with some fancy R graphics, nobody has really touched the trackdata yet.</p>",
      "rawMarkdown": "You cannot use R to handle that size of data, even loaded to Kaggle kernel.  And I find all of the public notebooks right now are simply playing with some fancy R graphics, nobody has really touched the trackdata yet.",
      "replies": [
        {
          "id": 697478,
          "postDate": "2019-12-18T01:27:37.960Z",
          "content": "<p>True</p>",
          "rawMarkdown": "True"
        }
      ]
    },
    {
      "id": 698160,
      "postDate": "2019-12-18T21:45:30.403Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 691741,
      "author_name": "Rasyid Ridha",
      "author_url": "",
      "post_date": "2019-12-10T13:02:43.073000",
      "content": "<p>Use Kaggle kernel dude <a href=\"/harnagpal\">@harnagpal</a>\nIt provides 16 GB RAM :) \n<a href=\"https://www.kaggle.com/docs/kernels\">https://www.kaggle.com/docs/kernels</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 691788,
          "author_name": "Kröger",
          "author_url": "",
          "post_date": "2019-12-10T13:54:19.393000",
          "content": "<p>I would suggest the same.\nIf you however insist on using your local machine then you could partially load the dataset through data.table library's function fread() playing with parameters skip and nrow.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 692727,
          "author_name": "Harish Nagpal",
          "author_url": "",
          "post_date": "2019-12-11T17:05:21.427000",
          "content": "<p>Thanks Rasyid, Kröger\nThat's great. I didn't know that we can use Kaggle like this.\nMeanwhile I tried with RSQLite pacakge. And it seems its working for me. I created a database and imported csv files into it in tables and then called them through data.table.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 694618,
      "author_name": "ReservePrism",
      "author_url": "",
      "post_date": "2019-12-13T21:10:59.677000",
      "content": "<p>You cannot use R to handle that size of data, even loaded to Kaggle kernel.  And I find all of the public notebooks right now are simply playing with some fancy R graphics, nobody has really touched the trackdata yet.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 697478,
          "author_name": "Harish Nagpal",
          "author_url": "",
          "post_date": "2019-12-18T01:27:37.960000",
          "content": "<p>True</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 698160,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-18T21:45:30.403000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "691741": "Use Kaggle kernel dude @harnagpal\nIt provides 16 GB RAM :) \nhttps://www.kaggle.com/docs/kernels",
    "691726": " All, any idea how to handle this large data set of 76 million records in R. I am using R and R hangs if I try to fetch data from this data set. \nThanks in advance.\nHowever if I do analysis in SQLite Studio, then it works although some what slow.\nI have 4 GB RAM on my PC.",
    "694618": "You cannot use R to handle that size of data, even loaded to Kaggle kernel.  And I find all of the public notebooks right now are simply playing with some fancy R graphics, nobody has really touched the trackdata yet.",
    "698160": ""
  }
}