{
  "id": 72287,
  "title": "Dataset Issues",
  "url": "/competitions/PLAsTiCC-2018/discussion/72287",
  "author_name": "",
  "post_date": "2018-11-22T01:33:23.965560Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>When i download the 7.45GB dataset and extract them, i am unable to fetch the full dataset; unable to fetch the dump. I wanted to check if the source dataset is corrupted and if so, where can i get the right datasset</p>",
  "messages": [
    {
      "id": "425698",
      "postDate": "11/22/2018 01:33:23",
      "content": "<p>When i download the 7.45GB dataset and extract them, i am unable to fetch the full dataset; unable to fetch the dump. I wanted to check if the source dataset is corrupted and if so, where can i get the right datasset</p>",
      "rawMarkdown": "When i download the 7.45GB dataset and extract them, i am unable to fetch the full dataset; unable to fetch the dump. I wanted to check if the source dataset is corrupted and if so, where can i get the right datasset",
      "votes": null
    },
    {
      "id": "426006",
      "postDate": "11/22/2018 12:33:38",
      "content": "<p>Is there a checksum for this dataset file? To verify if it is valid, you can use a checksum. If it is valid, you could verify if you have enough HD free space or, maybe, the software you are using to extract the dataset hasn't enough memory to do the job. Try using a command line to extract this dataset and reduce the overall memory footprint.</p>",
      "rawMarkdown": "Is there a checksum for this dataset file? To verify if it is valid, you can use a checksum. If it is valid, you could verify if you have enough HD free space or, maybe, the software you are using to extract the dataset hasn't enough memory to do the job. Try using a command line to extract this dataset and reduce the overall memory footprint.",
      "votes": null
    },
    {
      "id": "426010",
      "postDate": "11/22/2018 12:36:12",
      "content": "<p>There is a know issue with unziping zip files on mac os.  If you are on mac os then search the forum someone else had a similar issue and provided his solution.</p>",
      "rawMarkdown": "There is a know issue with unziping zip files on mac os.  If you are on mac os then search the forum someone else had a similar issue and provided his solution.",
      "votes": null
    },
    {
      "id": "426321",
      "postDate": "11/23/2018 04:36:47",
      "content": "<p>The dataset gave me fits in the beginning.  There are some good kernels out there.  ‘Fast test set reading’ is a really good one.  There are also kernels that read the file in chunks.</p>\n\n<p>I definitely advocate working in the Kaggle kernels.  Then you don’t have to download anything and you can access your work from anywhere that has a network connection.</p>",
      "rawMarkdown": "The dataset gave me fits in the beginning.  There are some good kernels out there.  ‘Fast test set reading’ is a really good one.  There are also kernels that read the file in chunks.\n\nI definitely advocate working in the Kaggle kernels.  Then you don’t have to download anything and you can access your work from anywhere that has a network connection.",
      "votes": null
    },
    {
      "id": "436055",
      "postDate": "12/09/2018 13:13:22",
      "content": "<p>Thanks; i was able to finally download the data-files and work on it. \nBut still a question though; is there some limits on using the compute processing capacity in Kaggle kernels?.  we had trouble to run featurisation using CesiumML on the test_set which was almost like 20GB and to finally derived the \"featured\" dataset. we had to transfer the dataset to an AWS EC2 cluster and then do the featurisation as our laptops were only of normal configuration.</p>",
      "rawMarkdown": "Thanks; i was able to finally download the data-files and work on it. \nBut still a question though; is there some limits on using the compute processing capacity in Kaggle kernels?.  we had trouble to run featurisation using CesiumML on the test_set which was almost like 20GB and to finally derived the \"featured\" dataset. we had to transfer the dataset to an AWS EC2 cluster and then do the featurisation as our laptops were only of normal configuration.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 426006,
      "author_name": "flaviocysne",
      "author_url": "",
      "post_date": "11/22/2018 12:33:38",
      "content": "<p>Is there a checksum for this dataset file? To verify if it is valid, you can use a checksum. If it is valid, you could verify if you have enough HD free space or, maybe, the software you are using to extract the dataset hasn't enough memory to do the job. Try using a command line to extract this dataset and reduce the overall memory footprint.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 426010,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "11/22/2018 12:36:12",
      "content": "<p>There is a know issue with unziping zip files on mac os.  If you are on mac os then search the forum someone else had a similar issue and provided his solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 426321,
      "author_name": "jimpsull",
      "author_url": "",
      "post_date": "11/23/2018 04:36:47",
      "content": "<p>The dataset gave me fits in the beginning.  There are some good kernels out there.  ‘Fast test set reading’ is a really good one.  There are also kernels that read the file in chunks.</p>\n\n<p>I definitely advocate working in the Kaggle kernels.  Then you don’t have to download anything and you can access your work from anywhere that has a network connection.</p>",
      "votes": null,
      "replies": [
        {
          "id": 436055,
          "author_name": "eashwar2018",
          "author_url": "",
          "post_date": "12/09/2018 13:13:22",
          "content": "<p>Thanks; i was able to finally download the data-files and work on it. \nBut still a question though; is there some limits on using the compute processing capacity in Kaggle kernels?.  we had trouble to run featurisation using CesiumML on the test_set which was almost like 20GB and to finally derived the \"featured\" dataset. we had to transfer the dataset to an AWS EC2 cluster and then do the featurisation as our laptops were only of normal configuration.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "425698": "When i download the 7.45GB dataset and extract them, i am unable to fetch the full dataset; unable to fetch the dump. I wanted to check if the source dataset is corrupted and if so, where can i get the right datasset",
    "426006": "Is there a checksum for this dataset file? To verify if it is valid, you can use a checksum. If it is valid, you could verify if you have enough HD free space or, maybe, the software you are using to extract the dataset hasn't enough memory to do the job. Try using a command line to extract this dataset and reduce the overall memory footprint.",
    "426010": "There is a know issue with unziping zip files on mac os.  If you are on mac os then search the forum someone else had a similar issue and provided his solution.",
    "426321": "The dataset gave me fits in the beginning.  There are some good kernels out there.  ‘Fast test set reading’ is a really good one.  There are also kernels that read the file in chunks.\n\nI definitely advocate working in the Kaggle kernels.  Then you don’t have to download anything and you can access your work from anywhere that has a network connection.",
    "436055": "Thanks; i was able to finally download the data-files and work on it. \nBut still a question though; is there some limits on using the compute processing capacity in Kaggle kernels?.  we had trouble to run featurisation using CesiumML on the test_set which was almost like 20GB and to finally derived the \"featured\" dataset. we had to transfer the dataset to an AWS EC2 cluster and then do the featurisation as our laptops were only of normal configuration."
  },
  "source": "meta"
}