{
  "id": 81679,
  "title": "How to read test.zip files on the cloud / submit with kernels?",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/81679",
  "author_name": "",
  "post_date": "2019-02-24T02:23:55.426109800Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi all,</p>\n\n<p>I am having trouble reading the test.zip file in the cloud to both to read the acoustic data and to predict the time to failure.</p>\n\n<p>The following code results in a file not found error:</p>\n\n<p>import zipfile\nwith zipfile.ZipFile('../input/test.zip') as z:\n    z.extractall('.')</p>\n\n<p>I have read previous threads about finding and extracting files...</p>\n\n<p><a href=\"https://www.kaggle.com/dansbecker/finding-your-files-in-kaggle-kernels\">https://www.kaggle.com/dansbecker/finding-your-files-in-kaggle-kernels</a>\n<a href=\"https://www.kaggle.com/zhikhe/how-to-read-datasets\">https://www.kaggle.com/zhikhe/how-to-read-datasets</a></p>\n\n<p>... and nothing seems to be working</p>\n\n<p>If anyone has advice about how to do this on the cloud that would be great.</p>\n\n<p>Please upvote if you have the same issue or question.</p>\n\n<p>Thanks!\n-Tom</p>",
  "messages": [
    {
      "id": "477160",
      "postDate": "02/24/2019 02:23:55",
      "content": "<p>Hi all,</p>\n\n<p>I am having trouble reading the test.zip file in the cloud to both to read the acoustic data and to predict the time to failure.</p>\n\n<p>The following code results in a file not found error:</p>\n\n<p>import zipfile\nwith zipfile.ZipFile('../input/test.zip') as z:\n    z.extractall('.')</p>\n\n<p>I have read previous threads about finding and extracting files...</p>\n\n<p><a href=\"https://www.kaggle.com/dansbecker/finding-your-files-in-kaggle-kernels\">https://www.kaggle.com/dansbecker/finding-your-files-in-kaggle-kernels</a>\n<a href=\"https://www.kaggle.com/zhikhe/how-to-read-datasets\">https://www.kaggle.com/zhikhe/how-to-read-datasets</a></p>\n\n<p>... and nothing seems to be working</p>\n\n<p>If anyone has advice about how to do this on the cloud that would be great.</p>\n\n<p>Please upvote if you have the same issue or question.</p>\n\n<p>Thanks!\n-Tom</p>",
      "rawMarkdown": "Hi all,\n\nI am having trouble reading the test.zip file in the cloud to both to read the acoustic data and to predict the time to failure.\n\nThe following code results in a file not found error:\n\nimport zipfile\nwith zipfile.ZipFile('../input/test.zip') as z:\n    z.extractall('.')\n\nI have read previous threads about finding and extracting files...\n\nhttps://www.kaggle.com/dansbecker/finding-your-files-in-kaggle-kernels\nhttps://www.kaggle.com/zhikhe/how-to-read-datasets\n\n... and nothing seems to be working\n\nIf anyone has advice about how to do this on the cloud that would be great.\n\nPlease upvote if you have the same issue or question.\n\nThanks!\n-Tom",
      "votes": null
    },
    {
      "id": "477454",
      "postDate": "02/24/2019 15:56:08",
      "content": "<p>Hey Tom! \nThe files are accessible directly without unzipping.</p>\n\n<p>I am using the glob library, and store all the tables in a list. You can try below code, it shall work : </p>\n\n<pre><code>test = []\nfor file in glob.glob(\"../input/test/*.csv\"):\n    test.append(file.split('/')[-1].split('.')[0])\ntest_df = []\nfor csv in test:\n    test_df.append(pd.read_csv(f'../input/test/{csv}.csv'))\n</code></pre>\n\n<p>(Probably not the more elegant way, but it works for what I want to do with it :) )</p>",
      "rawMarkdown": "Hey Tom! \nThe files are accessible directly without unzipping.\n\nI am using the glob library, and store all the tables in a list. You can try below code, it shall work : \n\n    test = []\n    for file in glob.glob(\"../input/test/*.csv\"):\n        test.append(file.split('/')[-1].split('.')[0])\n    test_df = []\n    for csv in test:\n        test_df.append(pd.read_csv(f'../input/test/{csv}.csv'))\n\n(Probably not the more elegant way, but it works for what I want to do with it :) )",
      "votes": null
    },
    {
      "id": "477514",
      "postDate": "02/24/2019 19:02:31",
      "content": "<p>Hi Jacky, thank you!  I just ran the code and \"test\" returns the segments and test_df returns the acoustic data.   I really appreciate it. </p>\n\n<p>I'll study the above code a bit more so I understand each part.  Thank you!</p>",
      "rawMarkdown": "Hi Jacky, thank you!  I just ran the code and \"test\" returns the segments and test_df returns the acoustic data.   I really appreciate it. \n\nI'll study the above code a bit more so I understand each part.  Thank you!",
      "votes": null
    },
    {
      "id": "477873",
      "postDate": "02/25/2019 12:05:24",
      "content": "<p>It worked. Thank you Jacky.</p>",
      "rawMarkdown": "It worked. Thank you Jacky.",
      "votes": null
    },
    {
      "id": "527634",
      "postDate": "05/06/2019 01:50:32",
      "content": "<p>Thanks Jacky.  It works well.</p>",
      "rawMarkdown": "Thanks Jacky.  It works well.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 477454,
      "author_name": "bowaka",
      "author_url": "",
      "post_date": "02/24/2019 15:56:08",
      "content": "<p>Hey Tom! \nThe files are accessible directly without unzipping.</p>\n\n<p>I am using the glob library, and store all the tables in a list. You can try below code, it shall work : </p>\n\n<pre><code>test = []\nfor file in glob.glob(\"../input/test/*.csv\"):\n    test.append(file.split('/')[-1].split('.')[0])\ntest_df = []\nfor csv in test:\n    test_df.append(pd.read_csv(f'../input/test/{csv}.csv'))\n</code></pre>\n\n<p>(Probably not the more elegant way, but it works for what I want to do with it :) )</p>",
      "votes": null,
      "replies": [
        {
          "id": 477514,
          "author_name": "tpmeli",
          "author_url": "",
          "post_date": "02/24/2019 19:02:31",
          "content": "<p>Hi Jacky, thank you!  I just ran the code and \"test\" returns the segments and test_df returns the acoustic data.   I really appreciate it. </p>\n\n<p>I'll study the above code a bit more so I understand each part.  Thank you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 477873,
          "author_name": "raosashankh",
          "author_url": "",
          "post_date": "02/25/2019 12:05:24",
          "content": "<p>It worked. Thank you Jacky.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 527634,
          "author_name": "yunlongzhang",
          "author_url": "",
          "post_date": "05/06/2019 01:50:32",
          "content": "<p>Thanks Jacky.  It works well.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "477160": "Hi all,\n\nI am having trouble reading the test.zip file in the cloud to both to read the acoustic data and to predict the time to failure.\n\nThe following code results in a file not found error:\n\nimport zipfile\nwith zipfile.ZipFile('../input/test.zip') as z:\n    z.extractall('.')\n\nI have read previous threads about finding and extracting files...\n\nhttps://www.kaggle.com/dansbecker/finding-your-files-in-kaggle-kernels\nhttps://www.kaggle.com/zhikhe/how-to-read-datasets\n\n... and nothing seems to be working\n\nIf anyone has advice about how to do this on the cloud that would be great.\n\nPlease upvote if you have the same issue or question.\n\nThanks!\n-Tom",
    "477454": "Hey Tom! \nThe files are accessible directly without unzipping.\n\nI am using the glob library, and store all the tables in a list. You can try below code, it shall work : \n\n    test = []\n    for file in glob.glob(\"../input/test/*.csv\"):\n        test.append(file.split('/')[-1].split('.')[0])\n    test_df = []\n    for csv in test:\n        test_df.append(pd.read_csv(f'../input/test/{csv}.csv'))\n\n(Probably not the more elegant way, but it works for what I want to do with it :) )",
    "477514": "Hi Jacky, thank you!  I just ran the code and \"test\" returns the segments and test_df returns the acoustic data.   I really appreciate it. \n\nI'll study the above code a bit more so I understand each part.  Thank you!",
    "477873": "It worked. Thank you Jacky.",
    "527634": "Thanks Jacky.  It works well."
  },
  "source": "meta"
}