{
  "id": 30746,
  "title": "is there a smaller download to avoid the 100 gig",
  "url": "/competitions/noaa-fisheries-steller-sea-lion-population-count/discussion/30746",
  "author_name": "AtomicOrbital",
  "post_date": "2017-03-27T22:10:56.349000",
  "votes": 13,
  "comment_count": 22,
  "views": 1,
  "content": "<p>Can someone post a much smaller input file ... say less than 100 meg ... for us to develop against then later when our code is ready we can download the full 100+ gig input file ?</p>\n\n<p>For those of us working from home our ISP will murder us if we download 100 gigs </p>",
  "messages": [
    {
      "id": 170882,
      "postDate": "2017-03-27T22:10:56.350Z",
      "content": "<p>Can someone post a much smaller input file ... say less than 100 meg ... for us to develop against then later when our code is ready we can download the full 100+ gig input file ?</p>\n\n<p>For those of us working from home our ISP will murder us if we download 100 gigs </p>",
      "rawMarkdown": "Can someone post a much smaller input file ... say less than 100 meg ... for us to develop against then later when our code is ready we can download the full 100+ gig input file ?\n\nFor those of us working from home our ISP will murder us if we download 100 gigs ",
      "votes": 13
    },
    {
      "id": 170897,
      "postDate": "2017-03-27T22:36:10.877Z",
      "content": "<p>Start an AWS instance ssh into it.  Download the files using wget and the cookie saved from your kaggle session.  Then dump the data to S3 for further analysis.  This can be free with the free tier AWS account</p>",
      "rawMarkdown": "Start an AWS instance ssh into it.  Download the files using wget and the cookie saved from your kaggle session.  Then dump the data to S3 for further analysis.  This can be free with the free tier AWS account",
      "votes": 7,
      "replies": [
        {
          "id": 177963,
          "postDate": "2017-04-26T09:46:30.700Z",
          "content": "<p>S3 storage for free tier is only 5GB, how can you dump 100GB data to S3 ?</p>",
          "rawMarkdown": "S3 storage for free tier is only 5GB, how can you dump 100GB data to S3 ?"
        }
      ]
    },
    {
      "id": 170896,
      "postDate": "2017-03-27T22:35:58.830Z",
      "content": "<p>I've put the new file — TrainSmall.7z — on the Data page.</p>",
      "rawMarkdown": "I've put the new file — TrainSmall.7z — on the Data page.",
      "votes": 4,
      "replies": [
        {
          "id": 171340,
          "postDate": "2017-03-29T15:17:19.670Z",
          "content": "<p>The TrainSmall.7z archive does not seem to follow the pattern of train and dotted images correctly. For example, see Train/3.jpg and TrainDotted/3.jpg. The same thing for 7, 9, etc.</p>",
          "rawMarkdown": "The TrainSmall.7z archive does not seem to follow the pattern of train and dotted images correctly. For example, see Train/3.jpg and TrainDotted/3.jpg. The same thing for 7, 9, etc.",
          "votes": 2
        }
      ]
    },
    {
      "id": 180973,
      "postDate": "2017-05-08T03:36:10.583Z",
      "content": "<p><a href=\"http://www.mediafire.com/file/qb2p0z94az91cgp/Train_Data.tar.gz\">http://www.mediafire.com/file/qb2p0z94az91cgp/Train_Data.tar.gz</a>\nI have made a new training dataset which is much easier to work with.</p>",
      "rawMarkdown": "[http://www.mediafire.com/file/qb2p0z94az91cgp/Train_Data.tar.gz][1]\nI have made a new training dataset which is much easier to work with.\n\n  [1]: http://www.mediafire.com/file/qb2p0z94az91cgp/Train_Data.tar.gz",
      "votes": 1
    },
    {
      "id": 171629,
      "postDate": "2017-03-30T21:01:10.183Z",
      "content": "<p>The training set discrepancies (described in <a href=\"https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/30895\">another thread</a>) affected the TrainSmall file disproportionately, so I've replaced it with TrainSmall2.7z, where all the images match the dotted versions.</p>",
      "rawMarkdown": "The training set discrepancies (described in [another thread](https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/30895)) affected the TrainSmall file disproportionately, so I've replaced it with TrainSmall2.7z, where all the images match the dotted versions.",
      "votes": 1
    },
    {
      "id": 170888,
      "postDate": "2017-03-27T22:19:54.523Z",
      "content": "<p>That's a good idea, thanks!</p>\n\n<p>Each image is ~5mb, so I could include the first 10 training images plus their dotted versions, along with the ground truth csv. Most images contain ~100 animals, so that should give you quite a few examples to start with. How does that sound?</p>",
      "rawMarkdown": "That's a good idea, thanks!\n\nEach image is ~5mb, so I could include the first 10 training images plus their dotted versions, along with the ground truth csv. Most images contain ~100 animals, so that should give you quite a few examples to start with. How does that sound?",
      "votes": 2,
      "replies": [
        {
          "id": 170892,
          "postDate": "2017-03-27T22:30:37.240Z",
          "content": "<p>nice ... that sounds more reasonable - Thanks</p>",
          "rawMarkdown": "nice ... that sounds more reasonable - Thanks"
        },
        {
          "id": 170894,
          "postDate": "2017-03-27T22:35:11.943Z",
          "content": "<p>Most people end up working with 512x512 or smaller images because of GPU constraints. I guess it would be more useful instead to have an alternative version of the dataset with smaller resolution. I don't have too much bandwidth available what I always end up doing is downloading the dataset in the cloud somewhere and resizing before downloading. Depending on the problem that could be 224x224 or even as low as 24x32 or similar :-)))</p>",
          "rawMarkdown": "Most people end up working with 512x512 or smaller images because of GPU constraints. I guess it would be more useful instead to have an alternative version of the dataset with smaller resolution. I don't have too much bandwidth available what I always end up doing is downloading the dataset in the cloud somewhere and resizing before downloading. Depending on the problem that could be 224x224 or even as low as 24x32 or similar :-)))",
          "votes": 3
        },
        {
          "id": 171208,
          "postDate": "2017-03-29T03:52:40.970Z",
          "content": "<p>I think for this one you need to split each image to few sub blocks, the resolution is already minimum\nfor finding the animals, especially the pup</p>",
          "rawMarkdown": "I think for this one you need to split each image to few sub blocks, the resolution is already minimum\nfor finding the animals, especially the pup",
          "votes": 1
        },
        {
          "id": 180974,
          "postDate": "2017-05-08T03:36:51.470Z",
          "content": "<p><a href=\"http://www.mediafire.com/file/qb2p0z94az91cgp/Train_Data.tar.gz\">Here is your solution</a></p>",
          "rawMarkdown": "[Here is your solution][1]\n\n\n  [1]: http://www.mediafire.com/file/qb2p0z94az91cgp/Train_Data.tar.gz",
          "votes": 1
        }
      ]
    },
    {
      "id": 175849,
      "postDate": "2017-04-17T20:45:45.737Z",
      "content": "<p>If someone who has already uncompressed the data can make a new torrent from that with each individual file in it, the rest of us could use our torrent managers to download only the subset of the data we have space for.</p>",
      "rawMarkdown": "If someone who has already uncompressed the data can make a new torrent from that with each individual file in it, the rest of us could use our torrent managers to download only the subset of the data we have space for."
    },
    {
      "id": 171085,
      "postDate": "2017-03-28T15:48:20.347Z",
      "content": "<p>Thanks for uploading a small data set.</p>\n\n<p>I think image 3.jpg without labels contains no animals. </p>",
      "rawMarkdown": "Thanks for uploading a small data set.\n\nI think image 3.jpg without labels contains no animals. ",
      "replies": [
        {
          "id": 171127,
          "postDate": "2017-03-28T18:56:15.933Z",
          "content": "<p>i found 1 at the bottom little bit to the right from the mid at the bottom</p>",
          "rawMarkdown": "i found 1 at the bottom little bit to the right from the mid at the bottom"
        },
        {
          "id": 171130,
          "postDate": "2017-03-28T19:04:03.197Z",
          "content": "<p>Hmm, thanks for flagging that, let me check on it.</p>",
          "rawMarkdown": "Hmm, thanks for flagging that, let me check on it."
        },
        {
          "id": 171289,
          "postDate": "2017-03-29T09:51:02.307Z",
          "content": "<p>I thought the images in folder train are exactly the same images in folder trainDotted but with dots!?</p>",
          "rawMarkdown": "I thought the images in folder train are exactly the same images in folder trainDotted but with dots!?"
        }
      ]
    },
    {
      "id": 171008,
      "postDate": "2017-03-28T10:28:24.740Z",
      "content": "<p>Good idea!</p>",
      "rawMarkdown": "Good idea!"
    },
    {
      "id": 170891,
      "postDate": "2017-03-27T22:27:17.327Z",
      "content": "<p>Having an area where we could grab a GB or less at a time would be very helpful.  I have to travel to a nearby town and use free wifi to avoid the $15.00/GB charge my phone carrier charges.  Free wifi is around 5-10mb throughput allowing 1 GB/hr download. </p>",
      "rawMarkdown": "Having an area where we could grab a GB or less at a time would be very helpful.  I have to travel to a nearby town and use free wifi to avoid the $15.00/GB charge my phone carrier charges.  Free wifi is around 5-10mb throughput allowing 1 GB/hr download. "
    },
    {
      "id": 170883,
      "postDate": "2017-03-27T22:12:23.690Z",
      "content": "<p>This could be very helpful for many of us ^^.</p>",
      "rawMarkdown": "This could be very helpful for many of us ^^."
    },
    {
      "id": 171561,
      "postDate": "2017-03-30T15:19:18.127Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 171630,
          "postDate": "2017-03-30T21:02:25.347Z",
          "content": "<p>Having coordinates would be great, but we have to work with the data we have. Removing the TrainDotted folder would save 5% of the filesize; most of the space is taken up by the test set.</p>",
          "rawMarkdown": "Having coordinates would be great, but we have to work with the data we have. Removing the TrainDotted folder would save 5% of the filesize; most of the space is taken up by the test set.",
          "votes": 1
        },
        {
          "id": 171679,
          "postDate": "2017-03-31T01:13:08.120Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 170897,
      "author_name": "Gareth Robins",
      "author_url": "",
      "post_date": "2017-03-27T22:36:10.877000",
      "content": "<p>Start an AWS instance ssh into it.  Download the files using wget and the cookie saved from your kaggle session.  Then dump the data to S3 for further analysis.  This can be free with the free tier AWS account</p>",
      "votes": 7,
      "replies": [
        {
          "id": 177963,
          "author_name": "RenzhiWu",
          "author_url": "",
          "post_date": "2017-04-26T09:46:30.700000",
          "content": "<p>S3 storage for free tier is only 5GB, how can you dump 100GB data to S3 ?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 170896,
      "author_name": "DataCanary",
      "author_url": "",
      "post_date": "2017-03-27T22:35:58.830000",
      "content": "<p>I've put the new file — TrainSmall.7z — on the Data page.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 171340,
          "author_name": "amaia",
          "author_url": "",
          "post_date": "2017-03-29T15:17:19.670000",
          "content": "<p>The TrainSmall.7z archive does not seem to follow the pattern of train and dotted images correctly. For example, see Train/3.jpg and TrainDotted/3.jpg. The same thing for 7, 9, etc.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 180973,
      "author_name": "Liam Larsen",
      "author_url": "",
      "post_date": "2017-05-08T03:36:10.583000",
      "content": "<p><a href=\"http://www.mediafire.com/file/qb2p0z94az91cgp/Train_Data.tar.gz\">http://www.mediafire.com/file/qb2p0z94az91cgp/Train_Data.tar.gz</a>\nI have made a new training dataset which is much easier to work with.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 171629,
      "author_name": "DataCanary",
      "author_url": "",
      "post_date": "2017-03-30T21:01:10.183000",
      "content": "<p>The training set discrepancies (described in <a href=\"https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/30895\">another thread</a>) affected the TrainSmall file disproportionately, so I've replaced it with TrainSmall2.7z, where all the images match the dotted versions.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 170888,
      "author_name": "DataCanary",
      "author_url": "",
      "post_date": "2017-03-27T22:19:54.523000",
      "content": "<p>That's a good idea, thanks!</p>\n\n<p>Each image is ~5mb, so I could include the first 10 training images plus their dotted versions, along with the ground truth csv. Most images contain ~100 animals, so that should give you quite a few examples to start with. How does that sound?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 170892,
          "author_name": "AtomicOrbital",
          "author_url": "",
          "post_date": "2017-03-27T22:30:37.240000",
          "content": "<p>nice ... that sounds more reasonable - Thanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 170894,
          "author_name": "amaia",
          "author_url": "",
          "post_date": "2017-03-27T22:35:11.943000",
          "content": "<p>Most people end up working with 512x512 or smaller images because of GPU constraints. I guess it would be more useful instead to have an alternative version of the dataset with smaller resolution. I don't have too much bandwidth available what I always end up doing is downloading the dataset in the cloud somewhere and resizing before downloading. Depending on the problem that could be 224x224 or even as low as 24x32 or similar :-)))</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 171208,
          "author_name": "zz",
          "author_url": "",
          "post_date": "2017-03-29T03:52:40.970000",
          "content": "<p>I think for this one you need to split each image to few sub blocks, the resolution is already minimum\nfor finding the animals, especially the pup</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 180974,
          "author_name": "Liam Larsen",
          "author_url": "",
          "post_date": "2017-05-08T03:36:51.470000",
          "content": "<p><a href=\"http://www.mediafire.com/file/qb2p0z94az91cgp/Train_Data.tar.gz\">Here is your solution</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 175849,
      "author_name": "Leif Poorman",
      "author_url": "",
      "post_date": "2017-04-17T20:45:45.737000",
      "content": "<p>If someone who has already uncompressed the data can make a new torrent from that with each individual file in it, the rest of us could use our torrent managers to download only the subset of the data we have space for.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 171085,
      "author_name": "Snowformatics",
      "author_url": "",
      "post_date": "2017-03-28T15:48:20.347000",
      "content": "<p>Thanks for uploading a small data set.</p>\n\n<p>I think image 3.jpg without labels contains no animals. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 171127,
          "author_name": "Chris Cross",
          "author_url": "",
          "post_date": "2017-03-28T18:56:15.933000",
          "content": "<p>i found 1 at the bottom little bit to the right from the mid at the bottom</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 171130,
          "author_name": "DataCanary",
          "author_url": "",
          "post_date": "2017-03-28T19:04:03.197000",
          "content": "<p>Hmm, thanks for flagging that, let me check on it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 171289,
          "author_name": "Snowformatics",
          "author_url": "",
          "post_date": "2017-03-29T09:51:02.307000",
          "content": "<p>I thought the images in folder train are exactly the same images in folder trainDotted but with dots!?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 171008,
      "author_name": "kaguser",
      "author_url": "",
      "post_date": "2017-03-28T10:28:24.740000",
      "content": "<p>Good idea!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 170891,
      "author_name": "Steve Armtrong",
      "author_url": "",
      "post_date": "2017-03-27T22:27:17.327000",
      "content": "<p>Having an area where we could grab a GB or less at a time would be very helpful.  I have to travel to a nearby town and use free wifi to avoid the $15.00/GB charge my phone carrier charges.  Free wifi is around 5-10mb throughput allowing 1 GB/hr download. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 170883,
      "author_name": "FernandoTN",
      "author_url": "",
      "post_date": "2017-03-27T22:12:23.690000",
      "content": "<p>This could be very helpful for many of us ^^.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 171561,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-03-30T15:19:18.127000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 171630,
          "author_name": "DataCanary",
          "author_url": "",
          "post_date": "2017-03-30T21:02:25.347000",
          "content": "<p>Having coordinates would be great, but we have to work with the data we have. Removing the TrainDotted folder would save 5% of the filesize; most of the space is taken up by the test set.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 171679,
          "author_name": "",
          "author_url": "",
          "post_date": "2017-03-31T01:13:08.120000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "170882": "Can someone post a much smaller input file ... say less than 100 meg ... for us to develop against then later when our code is ready we can download the full 100+ gig input file ?\n\nFor those of us working from home our ISP will murder us if we download 100 gigs ",
    "170897": "Start an AWS instance ssh into it.  Download the files using wget and the cookie saved from your kaggle session.  Then dump the data to S3 for further analysis.  This can be free with the free tier AWS account",
    "170896": "I've put the new file — TrainSmall.7z — on the Data page.",
    "180973": "[http://www.mediafire.com/file/qb2p0z94az91cgp/Train_Data.tar.gz][1]\nI have made a new training dataset which is much easier to work with.\n\n  [1]: http://www.mediafire.com/file/qb2p0z94az91cgp/Train_Data.tar.gz",
    "171629": "The training set discrepancies (described in [another thread](https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/30895)) affected the TrainSmall file disproportionately, so I've replaced it with TrainSmall2.7z, where all the images match the dotted versions.",
    "170888": "That's a good idea, thanks!\n\nEach image is ~5mb, so I could include the first 10 training images plus their dotted versions, along with the ground truth csv. Most images contain ~100 animals, so that should give you quite a few examples to start with. How does that sound?",
    "175849": "If someone who has already uncompressed the data can make a new torrent from that with each individual file in it, the rest of us could use our torrent managers to download only the subset of the data we have space for.",
    "171085": "Thanks for uploading a small data set.\n\nI think image 3.jpg without labels contains no animals. ",
    "171008": "Good idea!",
    "170891": "Having an area where we could grab a GB or less at a time would be very helpful.  I have to travel to a nearby town and use free wifi to avoid the $15.00/GB charge my phone carrier charges.  Free wifi is around 5-10mb throughput allowing 1 GB/hr download. ",
    "170883": "This could be very helpful for many of us ^^.",
    "171561": ""
  }
}