{
  "id": 26417,
  "title": "Is it possible to access that 30 GB (compressed file) without downloading it?",
  "url": "/competitions/outbrain-click-prediction/discussion/26417",
  "author_name": "",
  "post_date": "2016-12-13T06:23:53.310Z",
  "votes": null,
  "comment_count": 2,
  "views": 351,
  "content": "<p>I tried accessing the page_views.csv file by using Pandas.read_csv on the url (<a href=\"https://www.kaggle.com/c/outbrain-click-prediction/download/page_views.csv.zip\">https://www.kaggle.com/c/outbrain-click-prediction/download/page_views.csv.zip</a>)</p>\n\n<p>It doesn't seem to work. I added the read_csv option (compression='zip' and chunksize=1000) to handle the compression and enormous size of the file.</p>\n\n<p>Do I need to pass in login credentials or something? Or, is this a dead end.</p>\n\n<p>It would make sense if there were a way to remotely access the data rather than having every single Kaggler download their own copy.</p>\n\n<p>Thanks.\n-Tony</p>",
  "messages": [
    {
      "id": "149906",
      "postDate": "12/13/2016 06:23:53",
      "content": "<p>I tried accessing the page_views.csv file by using Pandas.read_csv on the url (<a href=\"https://www.kaggle.com/c/outbrain-click-prediction/download/page_views.csv.zip\">https://www.kaggle.com/c/outbrain-click-prediction/download/page_views.csv.zip</a>)</p>\n\n<p>It doesn't seem to work. I added the read_csv option (compression='zip' and chunksize=1000) to handle the compression and enormous size of the file.</p>\n\n<p>Do I need to pass in login credentials or something? Or, is this a dead end.</p>\n\n<p>It would make sense if there were a way to remotely access the data rather than having every single Kaggler download their own copy.</p>\n\n<p>Thanks.\n-Tony</p>",
      "rawMarkdown": "I tried accessing the page_views.csv file by using Pandas.read_csv on the url (https://www.kaggle.com/c/outbrain-click-prediction/download/page_views.csv.zip)\r\n\r\nIt doesn't seem to work. I added the read_csv option (compression='zip' and chunksize=1000) to handle the compression and enormous size of the file.\r\n\r\nDo I need to pass in login credentials or something? Or, is this a dead end.\r\n\r\nIt would make sense if there were a way to remotely access the data rather than having every single Kaggler download their own copy.\r\n\r\nThanks.\r\n-Tony",
      "votes": null
    },
    {
      "id": "149909",
      "postDate": "12/13/2016 06:30:05",
      "content": "<p>Well you can do that without needing to pass your credentials, if u check the Network tab under your browser console you will notice that your request to the above url is redirected to another url which does not require any credentials. Since the redirect is based on some cookie data (which is private to every account) I am not putting up that url. Then you can use the redirected url instead of this.</p>",
      "rawMarkdown": "Well you can do that without needing to pass your credentials, if u check the Network tab under your browser console you will notice that your request to the above url is redirected to another url which does not require any credentials. Since the redirect is based on some cookie data (which is private to every account) I am not putting up that url. Then you can use the redirected url instead of this.",
      "votes": null
    },
    {
      "id": "368167",
      "postDate": "08/09/2018 12:04:30",
      "content": "<p>I am trying to get a bigger sample of this dataset without importing. Any pointers about it?\nThanks in advance</p>",
      "rawMarkdown": "I am trying to get a bigger sample of this dataset without importing. Any pointers about it?\nThanks in advance",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 149909,
      "author_name": "keerath",
      "author_url": "",
      "post_date": "12/13/2016 06:30:05",
      "content": "<p>Well you can do that without needing to pass your credentials, if u check the Network tab under your browser console you will notice that your request to the above url is redirected to another url which does not require any credentials. Since the redirect is based on some cookie data (which is private to every account) I am not putting up that url. Then you can use the redirected url instead of this.</p>",
      "votes": null,
      "replies": [
        {
          "id": 368167,
          "author_name": "pragalbh1",
          "author_url": "",
          "post_date": "08/09/2018 12:04:30",
          "content": "<p>I am trying to get a bigger sample of this dataset without importing. Any pointers about it?\nThanks in advance</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "149906": "I tried accessing the page_views.csv file by using Pandas.read_csv on the url (https://www.kaggle.com/c/outbrain-click-prediction/download/page_views.csv.zip)\r\n\r\nIt doesn't seem to work. I added the read_csv option (compression='zip' and chunksize=1000) to handle the compression and enormous size of the file.\r\n\r\nDo I need to pass in login credentials or something? Or, is this a dead end.\r\n\r\nIt would make sense if there were a way to remotely access the data rather than having every single Kaggler download their own copy.\r\n\r\nThanks.\r\n-Tony",
    "149909": "Well you can do that without needing to pass your credentials, if u check the Network tab under your browser console you will notice that your request to the above url is redirected to another url which does not require any credentials. Since the redirect is based on some cookie data (which is private to every account) I am not putting up that url. Then you can use the redirected url instead of this.",
    "368167": "I am trying to get a bigger sample of this dataset without importing. Any pointers about it?\nThanks in advance"
  },
  "source": "meta"
}