{
  "id": 168439,
  "title": "How should TFRecords be formatted?",
  "url": "/competitions/landmark-retrieval-2020/discussion/168439",
  "author_name": "Fateh Aliyev",
  "post_date": "2020-07-20T16:23:16.431000",
  "votes": 0,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hi, I am trying to create a TFRecord and I am slightly confused on how they should be formatted.  Should they be images and then have probabilities for each number class (1-203K) or should they be formatted a different way?  I have tried looking at previous challenges involving TPUs but  I wasn't able to figure out how targets were presented?  Please Help!</p>",
  "messages": [
    {
      "id": 937037,
      "postDate": "2020-07-20T17:22:46.433Z",
      "content": "<p>Have you tried the following?</p>\n<p><a href=\"https://github.com/tensorflow/models/tree/master/research/delf/delf/python/training#prepare-the-data-for-training\" target=\"_blank\">https://github.com/tensorflow/models/tree/master/research/delf/delf/python/training#prepare-the-data-for-training</a></p>",
      "rawMarkdown": "Have you tried the following?\n\nhttps://github.com/tensorflow/models/tree/master/research/delf/delf/python/training#prepare-the-data-for-training",
      "votes": 1
    },
    {
      "id": 937587,
      "postDate": "2020-07-21T05:03:28.183Z",
      "content": "<p>Any luck with this?</p>",
      "rawMarkdown": "Any luck with this?",
      "votes": 2,
      "replies": [
        {
          "id": 938386,
          "postDate": "2020-07-21T13:48:17.570Z",
          "content": "<p>Well since the project is so big I have to go out and buy a hard drive to store it.  I think I should be able to get a TFRecord dataset out buy Thurs./Fri.</p>",
          "rawMarkdown": "Well since the project is so big I have to go out and buy a hard drive to store it.  I think I should be able to get a TFRecord dataset out buy Thurs./Fri.",
          "votes": 2
        },
        {
          "id": 938486,
          "postDate": "2020-07-21T15:09:30.323Z",
          "content": "<p>Are you trying to create tfrecords for the dataset available for download through the data section of the competition(~108 GB) or the official gldv2 dataset which is worth ~500 GB?</p>",
          "rawMarkdown": "Are you trying to create tfrecords for the dataset available for download through the data section of the competition(~108 GB) or the official gldv2 dataset which is worth ~500 GB?",
          "votes": 6
        },
        {
          "id": 938559,
          "postDate": "2020-07-21T15:52:42.923Z",
          "content": "<p>To be honest, any would work.  Is there a big difference between the 2?  Is the official gldv2 dataset just an extension?  I'm still fairly uncertain about the differences between the 2, such as the difference between train.csv and train_clean.csv?</p>",
          "rawMarkdown": "To be honest, any would work.  Is there a big difference between the 2?  Is the official gldv2 dataset just an extension?  I'm still fairly uncertain about the differences between the 2, such as the difference between train.csv and train_clean.csv?",
          "votes": 2
        },
        {
          "id": 938896,
          "postDate": "2020-07-21T20:50:51.397Z",
          "content": "<p>train_clean.csv does the job. It takes into account the current competition data(~108 GB) and creates tfrecords for that.</p>",
          "rawMarkdown": "train_clean.csv does the job. It takes into account the current competition data(~108 GB) and creates tfrecords for that.",
          "votes": 6
        }
      ]
    },
    {
      "id": 937256,
      "postDate": "2020-07-20T21:30:10.750Z",
      "content": "<p><a href=\"/pukkinming\">@pukkinming</a> it seems to me that the landmark ids are encoded as an _int64_feature but when I try using my own dataset encoded in the same way during training the loss equates to nan and the accuracy stays at 0. What am I missing?</p>",
      "rawMarkdown": "@pukkinming it seems to me that the landmark ids are encoded as an _int64_feature but when I try using my own dataset encoded in the same way during training the loss equates to nan and the accuracy stays at 0. What am I missing?"
    },
    {
      "id": 937123,
      "postDate": "2020-07-20T18:29:03.060Z",
      "content": "<p>Would I be able to simply plug in the dataset for the train directories?</p>",
      "rawMarkdown": "Would I be able to simply plug in the dataset for the train directories?"
    },
    {
      "id": 937082,
      "postDate": "2020-07-20T17:57:23.310Z",
      "content": "<p>Hi, slightly confused here, in build_tfrecords.py they use the dataset GLDv2. Is this the same thing as the dataset here? If not, how would I convert this to work with the dataset from this competition?</p>",
      "rawMarkdown": "Hi, slightly confused here, in build_tfrecords.py they use the dataset GLDv2. Is this the same thing as the dataset here? If not, how would I convert this to work with the dataset from this competition?",
      "replies": [
        {
          "id": 937113,
          "postDate": "2020-07-20T18:20:18.467Z",
          "content": "<p>I believe you can manipulate the parameters you pass to the following:</p>\n<p>python3 build<em>image</em>dataset.py \\<br>\n  --train<em>csv</em>path=gldv2<em>dataset/train/train.csv \\\n  --train</em>clean<em>csv</em>path=gldv2<em>dataset/train/train</em>clean.csv \\<br>\n  --train<em>directory=gldv2</em>dataset/train/<em>/</em>/*/ \\<br>\n  --output<em>directory=gldv2</em>dataset/tfrecord/ \\<br>\n  --num<em>shards=128 \\\n  --generate</em>train<em>validation</em>splits \\<br>\n  --validation<em>split</em>size=0.2</p>",
          "rawMarkdown": "I believe you can manipulate the parameters you pass to the following:\n\npython3 build_image_dataset.py \\\n  --train_csv_path=gldv2_dataset/train/train.csv \\\n  --train_clean_csv_path=gldv2_dataset/train/train_clean.csv \\\n  --train_directory=gldv2_dataset/train/*/*/*/ \\\n  --output_directory=gldv2_dataset/tfrecord/ \\\n  --num_shards=128 \\\n  --generate_train_validation_splits \\\n  --validation_split_size=0.2"
        },
        {
          "id": 937294,
          "postDate": "2020-07-20T22:52:15.057Z",
          "content": "<p>You can change the parameters accordingly. For example, you can replace the train.csv with the files you specify. </p>",
          "rawMarkdown": "You can change the parameters accordingly. For example, you can replace the train.csv with the files you specify. "
        }
      ]
    },
    {
      "id": 936971,
      "postDate": "2020-07-20T16:23:16.430Z",
      "content": "<p>Hi, I am trying to create a TFRecord and I am slightly confused on how they should be formatted.  Should they be images and then have probabilities for each number class (1-203K) or should they be formatted a different way?  I have tried looking at previous challenges involving TPUs but  I wasn't able to figure out how targets were presented?  Please Help!</p>",
      "rawMarkdown": "Hi, I am trying to create a TFRecord and I am slightly confused on how they should be formatted.  Should they be images and then have probabilities for each number class (1-203K) or should they be formatted a different way?  I have tried looking at previous challenges involving TPUs but  I wasn't able to figure out how targets were presented?  Please Help!"
    },
    {
      "id": 937081,
      "postDate": "2020-07-20T17:56:47.203Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 937037,
      "author_name": "FP",
      "author_url": "",
      "post_date": "2020-07-20T17:22:46.433000",
      "content": "<p>Have you tried the following?</p>\n<p><a href=\"https://github.com/tensorflow/models/tree/master/research/delf/delf/python/training#prepare-the-data-for-training\" target=\"_blank\">https://github.com/tensorflow/models/tree/master/research/delf/delf/python/training#prepare-the-data-for-training</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 937587,
      "author_name": "Chandan Verma",
      "author_url": "",
      "post_date": "2020-07-21T05:03:28.183000",
      "content": "<p>Any luck with this?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 938386,
          "author_name": "Fateh Aliyev",
          "author_url": "",
          "post_date": "2020-07-21T13:48:17.570000",
          "content": "<p>Well since the project is so big I have to go out and buy a hard drive to store it.  I think I should be able to get a TFRecord dataset out buy Thurs./Fri.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 938486,
          "author_name": "Chandan Verma",
          "author_url": "",
          "post_date": "2020-07-21T15:09:30.323000",
          "content": "<p>Are you trying to create tfrecords for the dataset available for download through the data section of the competition(~108 GB) or the official gldv2 dataset which is worth ~500 GB?</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 938559,
          "author_name": "Fateh Aliyev",
          "author_url": "",
          "post_date": "2020-07-21T15:52:42.923000",
          "content": "<p>To be honest, any would work.  Is there a big difference between the 2?  Is the official gldv2 dataset just an extension?  I'm still fairly uncertain about the differences between the 2, such as the difference between train.csv and train_clean.csv?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 938896,
          "author_name": "Chandan Verma",
          "author_url": "",
          "post_date": "2020-07-21T20:50:51.397000",
          "content": "<p>train_clean.csv does the job. It takes into account the current competition data(~108 GB) and creates tfrecords for that.</p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 937256,
      "author_name": "Fateh Aliyev",
      "author_url": "",
      "post_date": "2020-07-20T21:30:10.750000",
      "content": "<p><a href=\"/pukkinming\">@pukkinming</a> it seems to me that the landmark ids are encoded as an _int64_feature but when I try using my own dataset encoded in the same way during training the loss equates to nan and the accuracy stays at 0. What am I missing?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 937123,
      "author_name": "Fateh Aliyev",
      "author_url": "",
      "post_date": "2020-07-20T18:29:03.060000",
      "content": "<p>Would I be able to simply plug in the dataset for the train directories?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 937082,
      "author_name": "Fateh Aliyev",
      "author_url": "",
      "post_date": "2020-07-20T17:57:23.310000",
      "content": "<p>Hi, slightly confused here, in build_tfrecords.py they use the dataset GLDv2. Is this the same thing as the dataset here? If not, how would I convert this to work with the dataset from this competition?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 937113,
          "author_name": "FP",
          "author_url": "",
          "post_date": "2020-07-20T18:20:18.467000",
          "content": "<p>I believe you can manipulate the parameters you pass to the following:</p>\n<p>python3 build<em>image</em>dataset.py \\<br>\n  --train<em>csv</em>path=gldv2<em>dataset/train/train.csv \\\n  --train</em>clean<em>csv</em>path=gldv2<em>dataset/train/train</em>clean.csv \\<br>\n  --train<em>directory=gldv2</em>dataset/train/<em>/</em>/*/ \\<br>\n  --output<em>directory=gldv2</em>dataset/tfrecord/ \\<br>\n  --num<em>shards=128 \\\n  --generate</em>train<em>validation</em>splits \\<br>\n  --validation<em>split</em>size=0.2</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 937294,
          "author_name": "FP",
          "author_url": "",
          "post_date": "2020-07-20T22:52:15.057000",
          "content": "<p>You can change the parameters accordingly. For example, you can replace the train.csv with the files you specify. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 937081,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-20T17:56:47.203000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "937037": "Have you tried the following?\n\nhttps://github.com/tensorflow/models/tree/master/research/delf/delf/python/training#prepare-the-data-for-training",
    "937587": "Any luck with this?",
    "937256": "@pukkinming it seems to me that the landmark ids are encoded as an _int64_feature but when I try using my own dataset encoded in the same way during training the loss equates to nan and the accuracy stays at 0. What am I missing?",
    "937123": "Would I be able to simply plug in the dataset for the train directories?",
    "937082": "Hi, slightly confused here, in build_tfrecords.py they use the dataset GLDv2. Is this the same thing as the dataset here? If not, how would I convert this to work with the dataset from this competition?",
    "936971": "Hi, I am trying to create a TFRecord and I am slightly confused on how they should be formatted.  Should they be images and then have probabilities for each number class (1-203K) or should they be formatted a different way?  I have tried looking at previous challenges involving TPUs but  I wasn't able to figure out how targets were presented?  Please Help!",
    "937081": ""
  }
}