{
  "id": 177893,
  "title": "Complete TF-Record Dataset 90GB",
  "url": "/competitions/landmark-recognition-2020/discussion/177893",
  "author_name": "",
  "post_date": "2020-08-27T18:17:39.545857800Z",
  "votes": 17,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I have created TF Record Dataset : 90GB of Training Images.</p>\n<p><a href=\"https://www.kaggle.com/micheomaano/tfrecord-recognition-ds\" target=\"_blank\">https://www.kaggle.com/micheomaano/tfrecord-recognition-ds</a> is dataset link.</p>\n<p>I used the following serialization structure.<br>\nLabel_ID means Landmark_id, Label means label_encoder's output for Landmark_id</p>\n<p><code>def serialize_example(image, image_name, label_id, label):\n    feature = {}\n    feature['image'] = _bytes_feature(image)\n    feature['label_id'] = _int64_feature(label_id)\n    feature['label'] = _int64_feature(label)\n    example = tf.train.Example(features=tf.train.Features(feature=feature))\n    return example.SerializeToString()\n</code></p>\n<p>I used the following method to split Train and Validation:<br>\n<code>Thresh = 35\nxdf = df[df['counts'] &gt; Thresh]\ngrouped = xdf.groupby('landmark_id')\nvalid = grouped.apply(lambda x: x.sample(n = 10))\n</code><br>\nThis TF-Record Dataset is contains train split from above code.</p>\n<p>Let me know, if this helped.</p>",
  "messages": [
    {
      "id": "988047",
      "postDate": "08/27/2020 18:17:39",
      "content": "<p>I have created TF Record Dataset : 90GB of Training Images.</p>\n<p><a href=\"https://www.kaggle.com/micheomaano/tfrecord-recognition-ds\" target=\"_blank\">https://www.kaggle.com/micheomaano/tfrecord-recognition-ds</a> is dataset link.</p>\n<p>I used the following serialization structure.<br>\nLabel_ID means Landmark_id, Label means label_encoder's output for Landmark_id</p>\n<p><code>def serialize_example(image, image_name, label_id, label):\n    feature = {}\n    feature['image'] = _bytes_feature(image)\n    feature['label_id'] = _int64_feature(label_id)\n    feature['label'] = _int64_feature(label)\n    example = tf.train.Example(features=tf.train.Features(feature=feature))\n    return example.SerializeToString()\n</code></p>\n<p>I used the following method to split Train and Validation:<br>\n<code>Thresh = 35\nxdf = df[df['counts'] &gt; Thresh]\ngrouped = xdf.groupby('landmark_id')\nvalid = grouped.apply(lambda x: x.sample(n = 10))\n</code><br>\nThis TF-Record Dataset is contains train split from above code.</p>\n<p>Let me know, if this helped.</p>",
      "rawMarkdown": "I have created TF Record Dataset : 90GB of Training Images.\n\n\n\n[https://www.kaggle.com/micheomaano/tfrecord-recognition-ds](https://www.kaggle.com/micheomaano/tfrecord-recognition-ds) is dataset link.\n\nI used the following serialization structure.\nLabel_ID means Landmark_id, Label means label_encoder's output for Landmark_id\n\n`def serialize_example(image, image_name, label_id, label):\n    feature = {}\n    feature['image'] = _bytes_feature(image)\n    feature['label_id'] = _int64_feature(label_id)\n    feature['label'] = _int64_feature(label)\n    example = tf.train.Example(features=tf.train.Features(feature=feature))\n    return example.SerializeToString()\n`\n\nI used the following method to split Train and Validation:\n`Thresh = 35\nxdf = df[df['counts'] > Thresh]\ngrouped = xdf.groupby('landmark_id')\nvalid = grouped.apply(lambda x: x.sample(n = 10))\n`\nThis TF-Record Dataset is contains train split from above code.\n\nLet me know, if this helped.",
      "votes": null
    },
    {
      "id": "988056",
      "postDate": "08/27/2020 18:28:02",
      "content": "<p>Awesome! Thanks for this Salman</p>",
      "rawMarkdown": "Awesome! Thanks for this Salman",
      "votes": null
    },
    {
      "id": "988135",
      "postDate": "08/27/2020 20:04:25",
      "content": "<p>All Images are as it is. <br>\nNo one is resized.<br>\nI didn't performed any pre-processing.</p>",
      "rawMarkdown": "All Images are as it is. \nNo one is resized.\nI didn't performed any pre-processing.",
      "votes": null
    },
    {
      "id": "988441",
      "postDate": "08/28/2020 03:54:17",
      "content": "<p>Really amazing work salman!</p>",
      "rawMarkdown": "Really amazing work salman!",
      "votes": null
    },
    {
      "id": "997160",
      "postDate": "09/03/2020 19:40:35",
      "content": "<p>Thank you so much for this!<br>\nQues: Does the tfrecords contain both train and validation part?<br>\nIf not how can I get the validation part of the data?</p>",
      "rawMarkdown": "Thank you so much for this!\nQues: Does the tfrecords contain both train and validation part?\nIf not how can I get the validation part of the data?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 988056,
      "author_name": "sshikamaru",
      "author_url": "",
      "post_date": "08/27/2020 18:28:02",
      "content": "<p>Awesome! Thanks for this Salman</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 988135,
      "author_name": "micheomaano",
      "author_url": "",
      "post_date": "08/27/2020 20:04:25",
      "content": "<p>All Images are as it is. <br>\nNo one is resized.<br>\nI didn't performed any pre-processing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 988441,
      "author_name": "shivam17818",
      "author_url": "",
      "post_date": "08/28/2020 03:54:17",
      "content": "<p>Really amazing work salman!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 997160,
      "author_name": "rarun2596",
      "author_url": "",
      "post_date": "09/03/2020 19:40:35",
      "content": "<p>Thank you so much for this!<br>\nQues: Does the tfrecords contain both train and validation part?<br>\nIf not how can I get the validation part of the data?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "988047": "I have created TF Record Dataset : 90GB of Training Images.\n\n\n\n[https://www.kaggle.com/micheomaano/tfrecord-recognition-ds](https://www.kaggle.com/micheomaano/tfrecord-recognition-ds) is dataset link.\n\nI used the following serialization structure.\nLabel_ID means Landmark_id, Label means label_encoder's output for Landmark_id\n\n`def serialize_example(image, image_name, label_id, label):\n    feature = {}\n    feature['image'] = _bytes_feature(image)\n    feature['label_id'] = _int64_feature(label_id)\n    feature['label'] = _int64_feature(label)\n    example = tf.train.Example(features=tf.train.Features(feature=feature))\n    return example.SerializeToString()\n`\n\nI used the following method to split Train and Validation:\n`Thresh = 35\nxdf = df[df['counts'] > Thresh]\ngrouped = xdf.groupby('landmark_id')\nvalid = grouped.apply(lambda x: x.sample(n = 10))\n`\nThis TF-Record Dataset is contains train split from above code.\n\nLet me know, if this helped.",
    "988056": "Awesome! Thanks for this Salman",
    "988135": "All Images are as it is. \nNo one is resized.\nI didn't performed any pre-processing.",
    "988441": "Really amazing work salman!",
    "997160": "Thank you so much for this!\nQues: Does the tfrecords contain both train and validation part?\nIf not how can I get the validation part of the data?"
  },
  "source": "meta"
}