{
  "id": 106148,
  "title": "Can't find any relationship between tsv files content and competition ImageIDs ",
  "url": "/competitions/open-images-2019-object-detection/discussion/106148",
  "author_name": "",
  "post_date": "2019-08-28T12:44:40.031466900Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I found there are tar.gz files with the complete train dataset, but also .tsv files containing the original URLs to download the dataset\n<a href=\"https://github.com/cvdfoundation/open-images-dataset#download-full-dataset-with-google-storage-transfer\">https://github.com/cvdfoundation/open-images-dataset#download-full-dataset-with-google-storage-transfer</a></p>\n\n<p>It contains 18TB, I think there are more images than in our dataset of about 500MB+</p>\n\n<p>I'd like to use these tsvs to create a Google Cloud Storage bucket via  <a href=\"https://cloud.google.com/storage/transfer/create-manage-transfer-console\">https://cloud.google.com/storage/transfer/create-manage-transfer-console</a></p>\n\n<p>But reading these .tsv files, I didn't find any way to our ImageIDs.  Perhaps via md5checksums.</p>\n\n<p>I think I'm missing something</p>\n\n<p>Is there any way to have the relationsip ImageID--&gt;url ?</p>\n\n<p><a href=\"https://github.com/cvdfoundation/open-images-dataset/issues/29\">https://github.com/cvdfoundation/open-images-dataset/issues/29</a></p>\n\n<p>Thanks in advance.</p>\n\n<p>Cheers</p>",
  "messages": [
    {
      "id": "610107",
      "postDate": "08/28/2019 12:44:40",
      "content": "<p>I found there are tar.gz files with the complete train dataset, but also .tsv files containing the original URLs to download the dataset\n<a href=\"https://github.com/cvdfoundation/open-images-dataset#download-full-dataset-with-google-storage-transfer\">https://github.com/cvdfoundation/open-images-dataset#download-full-dataset-with-google-storage-transfer</a></p>\n\n<p>It contains 18TB, I think there are more images than in our dataset of about 500MB+</p>\n\n<p>I'd like to use these tsvs to create a Google Cloud Storage bucket via  <a href=\"https://cloud.google.com/storage/transfer/create-manage-transfer-console\">https://cloud.google.com/storage/transfer/create-manage-transfer-console</a></p>\n\n<p>But reading these .tsv files, I didn't find any way to our ImageIDs.  Perhaps via md5checksums.</p>\n\n<p>I think I'm missing something</p>\n\n<p>Is there any way to have the relationsip ImageID--&gt;url ?</p>\n\n<p><a href=\"https://github.com/cvdfoundation/open-images-dataset/issues/29\">https://github.com/cvdfoundation/open-images-dataset/issues/29</a></p>\n\n<p>Thanks in advance.</p>\n\n<p>Cheers</p>",
      "rawMarkdown": "I found there are tar.gz files with the complete train dataset, but also .tsv files containing the original URLs to download the dataset\nhttps://github.com/cvdfoundation/open-images-dataset#download-full-dataset-with-google-storage-transfer\n\nIt contains 18TB, I think there are more images than in our dataset of about 500MB+\n\nI'd like to use these tsvs to create a Google Cloud Storage bucket via  https://cloud.google.com/storage/transfer/create-manage-transfer-console\n\nBut reading these .tsv files, I didn't find any way to our ImageIDs.  Perhaps via md5checksums.\n\nI think I'm missing something\n\nIs there any way to have the relationsip ImageID--&gt;url ?\n\nhttps://github.com/cvdfoundation/open-images-dataset/issues/29\n\nThanks in advance.\n\nCheers",
      "votes": null
    },
    {
      "id": "610354",
      "postDate": "08/28/2019 17:32:48",
      "content": "<p>I downloaded a sample image, and I calculated its md5sum:</p>\n\n<p>```\naws s3 --no-sign-request cp s3://open-images-dataset/train/0004009be735ce46.jpg ./\nopenssl dgst -md5 -binary 0004009be735ce46.jpg  | openssl enc -base64</p>\n\n<p>J4zLKvzlbHJhpReiAkNqcA==\n```</p>\n\n<p>But I didn't find this hash the .tsv files.  Are those .tsv files usefull at all?</p>",
      "rawMarkdown": "I downloaded a sample image, and I calculated its md5sum:\n\n```\naws s3 --no-sign-request cp s3://open-images-dataset/train/0004009be735ce46.jpg ./\nopenssl dgst -md5 -binary 0004009be735ce46.jpg  | openssl enc -base64\n\nJ4zLKvzlbHJhpReiAkNqcA==\n```\n\nBut I didn't find this hash the .tsv files.  Are those .tsv files usefull at all?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 610354,
      "author_name": "virilo",
      "author_url": "",
      "post_date": "08/28/2019 17:32:48",
      "content": "<p>I downloaded a sample image, and I calculated its md5sum:</p>\n\n<p>```\naws s3 --no-sign-request cp s3://open-images-dataset/train/0004009be735ce46.jpg ./\nopenssl dgst -md5 -binary 0004009be735ce46.jpg  | openssl enc -base64</p>\n\n<p>J4zLKvzlbHJhpReiAkNqcA==\n```</p>\n\n<p>But I didn't find this hash the .tsv files.  Are those .tsv files usefull at all?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "610107": "I found there are tar.gz files with the complete train dataset, but also .tsv files containing the original URLs to download the dataset\nhttps://github.com/cvdfoundation/open-images-dataset#download-full-dataset-with-google-storage-transfer\n\nIt contains 18TB, I think there are more images than in our dataset of about 500MB+\n\nI'd like to use these tsvs to create a Google Cloud Storage bucket via  https://cloud.google.com/storage/transfer/create-manage-transfer-console\n\nBut reading these .tsv files, I didn't find any way to our ImageIDs.  Perhaps via md5checksums.\n\nI think I'm missing something\n\nIs there any way to have the relationsip ImageID--&gt;url ?\n\nhttps://github.com/cvdfoundation/open-images-dataset/issues/29\n\nThanks in advance.\n\nCheers",
    "610354": "I downloaded a sample image, and I calculated its md5sum:\n\n```\naws s3 --no-sign-request cp s3://open-images-dataset/train/0004009be735ce46.jpg ./\nopenssl dgst -md5 -binary 0004009be735ce46.jpg  | openssl enc -base64\n\nJ4zLKvzlbHJhpReiAkNqcA==\n```\n\nBut I didn't find this hash the .tsv files.  Are those .tsv files usefull at all?"
  },
  "source": "meta"
}