{
  "id": 20180,
  "title": "Access File ?",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20180",
  "author_name": "",
  "post_date": "2016-04-16T18:50:02.110Z",
  "votes": null,
  "comment_count": 4,
  "views": 680,
  "content": "<p>Hi fellow kagglers since I am newbie would like to know how are you viewing data? My excel is not opening the training file as well as microsoft access says the file is to large. Please guide on how to go about this? how to view the data? I guess there must be some cloud based solution. Can someone help highlighting that.</p>\n\n<p>Thanks</p>",
  "messages": [
    {
      "id": "115158",
      "postDate": "04/16/2016 18:50:02",
      "content": "<p>Hi fellow kagglers since I am newbie would like to know how are you viewing data? My excel is not opening the training file as well as microsoft access says the file is to large. Please guide on how to go about this? how to view the data? I guess there must be some cloud based solution. Can someone help highlighting that.</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Hi fellow kagglers since I am newbie would like to know how are you viewing data? My excel is not opening the training file as well as microsoft access says the file is to large. Please guide on how to go about this? how to view the data? I guess there must be some cloud based solution. Can someone help highlighting that.\r\n\r\nThanks",
      "votes": null
    },
    {
      "id": "115160",
      "postDate": "04/16/2016 19:00:04",
      "content": "<p>You could use the head command in linux to split the file for first inspections. Then you could use spreadsheet or editor.</p>\n\n<p>I suspect we might need to either split data or process it sequentially. </p>\n\n<p>Because even if we are able to load train... We need more memory for test, destinations and not to mention the algorithms later.</p>\n\n<p>Cheers</p>\n\n<p>Gerhard</p>",
      "rawMarkdown": "You could use the head command in linux to split the file for first inspections. Then you could use spreadsheet or editor.\r\n\r\nI suspect we might need to either split data or process it sequentially. \r\n\r\nBecause even if we are able to load train... We need more memory for test, destinations and not to mention the algorithms later.\r\n\r\nCheers\r\n\r\nGerhard",
      "votes": null
    },
    {
      "id": "115168",
      "postDate": "04/16/2016 19:54:05",
      "content": "<p>This is not meant for Excel, especially the training data. If you use something like R or Python (with Pandas), you can read in a subset (say, the first million rows) or read it in chunks.</p>\n\n<p>So far I've spent most of my time reading up on this stuff. Here's what I was reading until a couple of minute ago:</p>\n\n<p><a href=\"https://www.reddit.com/r/MachineLearning/comments/2s1e8z/pandas_tears_large_data_set_and_memory_issues/\">https://www.reddit.com/r/MachineLearning/comments/2s1e8z/pandas_tears_large_data_set_and_memory_issues/</a></p>\n\n<p><a href=\"http://stackoverflow.com/questions/14262433/large-data-work-flows-using-pandas\">http://stackoverflow.com/questions/14262433/large-data-work-flows-using-pandas</a></p>",
      "rawMarkdown": "This is not meant for Excel, especially the training data. If you use something like R or Python (with Pandas), you can read in a subset (say, the first million rows) or read it in chunks.\r\n\r\nSo far I've spent most of my time reading up on this stuff. Here's what I was reading until a couple of minute ago:\r\n\r\nhttps://www.reddit.com/r/MachineLearning/comments/2s1e8z/pandas_tears_large_data_set_and_memory_issues/\r\n\r\nhttp://stackoverflow.com/questions/14262433/large-data-work-flows-using-pandas",
      "votes": null
    },
    {
      "id": "115230",
      "postDate": "04/17/2016 10:33:09",
      "content": "<p>If you are using R and just want to get a feel of  the training data you can use the parameter &quot;nrows = some small number&quot;  in your read.csv() call to read that many lines from the CSV file.  The test file opens in a text editor like Notepad++. If you have a Linux environment (I don't) you can try head/tail commands as Gerhard has suggested above. </p>\n\n<p>Cheers, </p>\n\n<p>Partha</p>",
      "rawMarkdown": "If you are using R and just want to get a feel of  the training data you can use the parameter \"nrows = some small number\"  in your read.csv() call to read that many lines from the CSV file.  The test file opens in a text editor like Notepad++. If you have a Linux environment (I don't) you can try head/tail commands as Gerhard has suggested above. \r\n\r\nCheers, \r\n\r\nPartha",
      "votes": null
    },
    {
      "id": "120001",
      "postDate": "05/14/2016 14:39:20",
      "content": "<p>[quote=navneethc;115168]</p>\n\n<p>This is not meant for Excel, especially the training data. If you use something like R or Python (with Pandas), you can read in a subset (say, the first million rows) or read it in chunks.</p>\n\n<p>So far I've spent most of my time reading up on this stuff. Here's what I was reading until a couple of minute ago:</p>\n\n<p><a href=\"https://www.reddit.com/r/MachineLearning/comments/2s1e8z/pandas_tears_large_data_set_and_memory_issues/\">https://www.reddit.com/r/MachineLearning/comments/2s1e8z/pandas_tears_large_data_set_and_memory_issues/</a></p>\n\n<p><a href=\"http://stackoverflow.com/questions/14262433/large-data-work-flows-using-pandas\">http://stackoverflow.com/questions/14262433/large-data-work-flows-using-pandas</a></p>\n\n<p>[/quote]</p>\n\n<p>good advice</p>",
      "rawMarkdown": "[quote=navneethc;115168]\r\n\r\nThis is not meant for Excel, especially the training data. If you use something like R or Python (with Pandas), you can read in a subset (say, the first million rows) or read it in chunks.\r\n\r\nSo far I've spent most of my time reading up on this stuff. Here's what I was reading until a couple of minute ago:\r\n\r\nhttps://www.reddit.com/r/MachineLearning/comments/2s1e8z/pandas_tears_large_data_set_and_memory_issues/\r\n\r\nhttp://stackoverflow.com/questions/14262433/large-data-work-flows-using-pandas\r\n\r\n[/quote]\r\n\r\ngood advice",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 115160,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "04/16/2016 19:00:04",
      "content": "<p>You could use the head command in linux to split the file for first inspections. Then you could use spreadsheet or editor.</p>\n\n<p>I suspect we might need to either split data or process it sequentially. </p>\n\n<p>Because even if we are able to load train... We need more memory for test, destinations and not to mention the algorithms later.</p>\n\n<p>Cheers</p>\n\n<p>Gerhard</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115168,
      "author_name": "navneethc",
      "author_url": "",
      "post_date": "04/16/2016 19:54:05",
      "content": "<p>This is not meant for Excel, especially the training data. If you use something like R or Python (with Pandas), you can read in a subset (say, the first million rows) or read it in chunks.</p>\n\n<p>So far I've spent most of my time reading up on this stuff. Here's what I was reading until a couple of minute ago:</p>\n\n<p><a href=\"https://www.reddit.com/r/MachineLearning/comments/2s1e8z/pandas_tears_large_data_set_and_memory_issues/\">https://www.reddit.com/r/MachineLearning/comments/2s1e8z/pandas_tears_large_data_set_and_memory_issues/</a></p>\n\n<p><a href=\"http://stackoverflow.com/questions/14262433/large-data-work-flows-using-pandas\">http://stackoverflow.com/questions/14262433/large-data-work-flows-using-pandas</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115230,
      "author_name": "pachakra",
      "author_url": "",
      "post_date": "04/17/2016 10:33:09",
      "content": "<p>If you are using R and just want to get a feel of  the training data you can use the parameter &quot;nrows = some small number&quot;  in your read.csv() call to read that many lines from the CSV file.  The test file opens in a text editor like Notepad++. If you have a Linux environment (I don't) you can try head/tail commands as Gerhard has suggested above. </p>\n\n<p>Cheers, </p>\n\n<p>Partha</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120001,
      "author_name": "bennie",
      "author_url": "",
      "post_date": "05/14/2016 14:39:20",
      "content": "<p>[quote=navneethc;115168]</p>\n\n<p>This is not meant for Excel, especially the training data. If you use something like R or Python (with Pandas), you can read in a subset (say, the first million rows) or read it in chunks.</p>\n\n<p>So far I've spent most of my time reading up on this stuff. Here's what I was reading until a couple of minute ago:</p>\n\n<p><a href=\"https://www.reddit.com/r/MachineLearning/comments/2s1e8z/pandas_tears_large_data_set_and_memory_issues/\">https://www.reddit.com/r/MachineLearning/comments/2s1e8z/pandas_tears_large_data_set_and_memory_issues/</a></p>\n\n<p><a href=\"http://stackoverflow.com/questions/14262433/large-data-work-flows-using-pandas\">http://stackoverflow.com/questions/14262433/large-data-work-flows-using-pandas</a></p>\n\n<p>[/quote]</p>\n\n<p>good advice</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "115158": "Hi fellow kagglers since I am newbie would like to know how are you viewing data? My excel is not opening the training file as well as microsoft access says the file is to large. Please guide on how to go about this? how to view the data? I guess there must be some cloud based solution. Can someone help highlighting that.\r\n\r\nThanks",
    "115160": "You could use the head command in linux to split the file for first inspections. Then you could use spreadsheet or editor.\r\n\r\nI suspect we might need to either split data or process it sequentially. \r\n\r\nBecause even if we are able to load train... We need more memory for test, destinations and not to mention the algorithms later.\r\n\r\nCheers\r\n\r\nGerhard",
    "115168": "This is not meant for Excel, especially the training data. If you use something like R or Python (with Pandas), you can read in a subset (say, the first million rows) or read it in chunks.\r\n\r\nSo far I've spent most of my time reading up on this stuff. Here's what I was reading until a couple of minute ago:\r\n\r\nhttps://www.reddit.com/r/MachineLearning/comments/2s1e8z/pandas_tears_large_data_set_and_memory_issues/\r\n\r\nhttp://stackoverflow.com/questions/14262433/large-data-work-flows-using-pandas",
    "115230": "If you are using R and just want to get a feel of  the training data you can use the parameter \"nrows = some small number\"  in your read.csv() call to read that many lines from the CSV file.  The test file opens in a text editor like Notepad++. If you have a Linux environment (I don't) you can try head/tail commands as Gerhard has suggested above. \r\n\r\nCheers, \r\n\r\nPartha",
    "120001": "[quote=navneethc;115168]\r\n\r\nThis is not meant for Excel, especially the training data. If you use something like R or Python (with Pandas), you can read in a subset (say, the first million rows) or read it in chunks.\r\n\r\nSo far I've spent most of my time reading up on this stuff. Here's what I was reading until a couple of minute ago:\r\n\r\nhttps://www.reddit.com/r/MachineLearning/comments/2s1e8z/pandas_tears_large_data_set_and_memory_issues/\r\n\r\nhttp://stackoverflow.com/questions/14262433/large-data-work-flows-using-pandas\r\n\r\n[/quote]\r\n\r\ngood advice"
  },
  "source": "meta"
}