{
  "id": 145205,
  "title": "How can I generate a dataset of pairs of images of length of the original training set?",
  "url": "/competitions/flower-classification-with-tpus/discussion/145205",
  "author_name": "",
  "post_date": "2020-04-22T09:11:51.466952200Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>So essentially the dataset will be twice the size of the original training set (same length, two columns)</p>\n\n<p>I have an algorithm to generate the pairs how I want them (in the sense of distribution of classes) but I am having trouble fitting it all in memory.</p>\n\n<p>Currently I am loading in the data as a dataset according to the starter notebook.</p>\n\n<p>But to generate the pairs I am iterating over this dataset to split it into a list of images and a list of the corresponding labels.</p>\n\n<p>then I am using these two generated lists to generate a list of pairs of images and list of labels signifiying if the pairs are of the same class or not.  </p>\n\n<p>THEN I am storing them in a dictionary and converting them to tensors </p>\n\n<p>At the end of this process I am already using 6gbs of ram just for the training set. </p>",
  "messages": [
    {
      "id": "816319",
      "postDate": "04/22/2020 09:11:51",
      "content": "<p>So essentially the dataset will be twice the size of the original training set (same length, two columns)</p>\n\n<p>I have an algorithm to generate the pairs how I want them (in the sense of distribution of classes) but I am having trouble fitting it all in memory.</p>\n\n<p>Currently I am loading in the data as a dataset according to the starter notebook.</p>\n\n<p>But to generate the pairs I am iterating over this dataset to split it into a list of images and a list of the corresponding labels.</p>\n\n<p>then I am using these two generated lists to generate a list of pairs of images and list of labels signifiying if the pairs are of the same class or not.  </p>\n\n<p>THEN I am storing them in a dictionary and converting them to tensors </p>\n\n<p>At the end of this process I am already using 6gbs of ram just for the training set. </p>",
      "rawMarkdown": "So essentially the dataset will be twice the size of the original training set (same length, two columns)\n\nI have an algorithm to generate the pairs how I want them (in the sense of distribution of classes) but I am having trouble fitting it all in memory.\n\nCurrently I am loading in the data as a dataset according to the starter notebook.\n\nBut to generate the pairs I am iterating over this dataset to split it into a list of images and a list of the corresponding labels.\n\nthen I am using these two generated lists to generate a list of pairs of images and list of labels signifiying if the pairs are of the same class or not.  \n\nTHEN I am storing them in a dictionary and converting them to tensors \n\nAt the end of this process I am already using 6gbs of ram just for the training set.",
      "votes": null
    },
    {
      "id": "816679",
      "postDate": "04/22/2020 14:05:28",
      "content": "<p>Try these out.</p>\n\n<p><a href=\"https://www.tensorflow.org/api_docs/python/tf/data/experimental/choose_from_datasets\">https://www.tensorflow.org/api_docs/python/tf/data/experimental/choose_from_datasets</a></p>\n\n<p>or</p>\n\n<p><a href=\"https://www.tensorflow.org/api_docs/python/tf/data/experimental/sample_from_datasets\">https://www.tensorflow.org/api_docs/python/tf/data/experimental/sample_from_datasets</a></p>\n\n<p>or just tf.data.Dataset.zip</p>\n\n<p>I put together some quick working examples of each of these in <a href=\"https://www.kaggle.com/calebeverett/combining-dataset-examples\">combining-dataset-examples</a>.</p>",
      "rawMarkdown": "Try these out.\n\nhttps://www.tensorflow.org/api_docs/python/tf/data/experimental/choose_from_datasets\n\nor\n\nhttps://www.tensorflow.org/api_docs/python/tf/data/experimental/sample_from_datasets\n\nor just tf.data.Dataset.zip\n\nI put together some quick working examples of each of these in [combining-dataset-examples](https://www.kaggle.com/calebeverett/combining-dataset-examples).",
      "votes": null
    },
    {
      "id": "817059",
      "postDate": "04/22/2020 19:55:41",
      "content": "<p>thank you for the links Ive been looking through them right now, but I have a up couple follow up questions if you don't mind.</p>\n\n<p>So this is the algorithm im using to generate the pairs, this is a barely tweaked version from some tutorial \n```\ndef make_pairs(x,y):\n    num_classes = max(y) + 1\n    digit_indices = [np.where(y == i)[0] for i in range(num_classes)]</p>\n\n<h1>pairs = []</h1>\n\n<pre><code>left_images = []\nright_images = []\nlabels = []\n\nfor idx1 in range(len(x)):\n    #add matching example\n    x1 = x[idx1]\n    label1 = y[idx1]\n    idx2 = random.choice(digit_indices[label1])\n    x2 = x[idx2]\n</code></pre>\n\n<h1>pairs+= [[x1, x2]]</h1>\n\n<pre><code>    left_images.append(x1)\n    right_images.append(x2)\n    labels += [1]\n\n    # add a not matching example\n    label2 = random.randint(0, num_classes-1)\n    while label2 == label1:\n        label2 = random.randint(0, num_classes-1)\n\n    idx2 = random.choice(digit_indices[label2])\n    x2 = x[idx2]\n</code></pre>\n\n<h1>pairs += [[x1,x2]]</h1>\n\n<pre><code>    left_images.append(x1)\n    right_images.append(x2)\n    labels += [0]\n\nreturn left_images, right_images, labels\n</code></pre>\n\n<p><code>\n</code>\ndef get_train(image_size):\n    images, labels = make_single_dataset(image_size, TRAINING_FILENAMES,labeled=True)\n    left_imgs, right_imgs, labels = make_pairs(images, labels)    </p>\n\n<pre><code>del images\ntrain_dict = {'input_1':left_imgs,\n              'input_2': right_imgs}\n\nout_dict = {'output_1':labels}\n\ndel left_imgs, right_imgs, labels\nprint('dictionaries generated')\ntrain_dict =  tf.data.Dataset.from_tensor_slices(train_dict)\nout_dict =  tf.data.Dataset.from_tensor_slices(out_dict)\n\nprint('tensors generated')\n\ntrain_dict = tf.data.Dataset.zip((train_dict, out_dict))\n\ndel out_dict\nprint('tensors combined')\ntrain_dict = train_dict.repeat()\ntrain_dict = train_dict.batch(BATCH_SIZE)\ntrain_dict = train_dict.prefetch(AUTO) # prefetch next batch while training (autotune prefetch buffer size)\n\n\n\n\n\nreturn train_dict\n</code></pre>\n\n<p>```</p>\n\n<p>So I generate a list of all images and labels by iterating over a dataset that the files were loaded into.  I pass those to a make_pairs function that goes through every image and gets another image of the same class and an image of the different class  class.  these are stored in three lists (left images, right images and labels). I store these into dictionaries with the corresponding label for the model then load them into tensors.    Later I zip these two tensors using data.zip</p>\n\n<p>Do you think I can leverage the functions you mentioned to make this process more efficient ?  By the time I'm finished loading the training set's input and output dictionaries I'm using almost 10gbs of ram.  I'm deleting the unused objects (correctly, I hope) but that only helped so much </p>",
      "rawMarkdown": "thank you for the links Ive been looking through them right now, but I have a up couple follow up questions if you don't mind.\n\nSo this is the algorithm im using to generate the pairs, this is a barely tweaked version from some tutorial \n```\ndef make_pairs(x,y):\n    num_classes = max(y) + 1\n    digit_indices = [np.where(y == i)[0] for i in range(num_classes)]\n#     pairs = []\n    left_images = []\n    right_images = []\n    labels = []\n    \n    for idx1 in range(len(x)):\n        #add matching example\n        x1 = x[idx1]\n        label1 = y[idx1]\n        idx2 = random.choice(digit_indices[label1])\n        x2 = x[idx2]\n        \n#         pairs+= [[x1, x2]]\n        left_images.append(x1)\n        right_images.append(x2)\n        labels += [1]\n        \n        # add a not matching example\n        label2 = random.randint(0, num_classes-1)\n        while label2 == label1:\n            label2 = random.randint(0, num_classes-1)\n            \n        idx2 = random.choice(digit_indices[label2])\n        x2 = x[idx2]\n        \n#         pairs += [[x1,x2]]\n        left_images.append(x1)\n        right_images.append(x2)\n        labels += [0]\n        \n    return left_images, right_images, labels\n```\n```\ndef get_train(image_size):\n    images, labels = make_single_dataset(image_size, TRAINING_FILENAMES,labeled=True)\n    left_imgs, right_imgs, labels = make_pairs(images, labels)    \n    \n    del images\n    train_dict = {'input_1':left_imgs,\n                  'input_2': right_imgs}\n\n    out_dict = {'output_1':labels}\n    \n    del left_imgs, right_imgs, labels\n    print('dictionaries generated')\n    train_dict =  tf.data.Dataset.from_tensor_slices(train_dict)\n    out_dict =  tf.data.Dataset.from_tensor_slices(out_dict)\n\n    print('tensors generated')\n\n    train_dict = tf.data.Dataset.zip((train_dict, out_dict))\n    \n    del out_dict\n    print('tensors combined')\n    train_dict = train_dict.repeat()\n    train_dict = train_dict.batch(BATCH_SIZE)\n    train_dict = train_dict.prefetch(AUTO) # prefetch next batch while training (autotune prefetch buffer size)\n\n\n\n\n  \n    return train_dict\n```\n\n\nSo I generate a list of all images and labels by iterating over a dataset that the files were loaded into.  I pass those to a make_pairs function that goes through every image and gets another image of the same class and an image of the different class  class.  these are stored in three lists (left images, right images and labels). I store these into dictionaries with the corresponding label for the model then load them into tensors.    Later I zip these two tensors using data.zip\n\nDo you think I can leverage the functions you mentioned to make this process more efficient ?  By the time I'm finished loading the training set's input and output dictionaries I'm using almost 10gbs of ram.  I'm deleting the unused objects (correctly, I hope) but that only helped so much",
      "votes": null
    },
    {
      "id": "817069",
      "postDate": "04/22/2020 20:06:57",
      "content": "<p>another quick thought I had is if there simply isn't enough memory to do this with the whole dataset, I can possibly use sample by dataset for n amount of images where n &lt; Len(training_files). by calculating proper class weights I may even be able to generate a balanced , albeit , smaller, dataset. </p>",
      "rawMarkdown": "another quick thought I had is if there simply isn't enough memory to do this with the whole dataset, I can possibly use sample by dataset for n amount of images where n &lt; Len(training_files). by calculating proper class weights I may even be able to generate a balanced , albeit , smaller, dataset.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 816679,
      "author_name": "calebeverett",
      "author_url": "",
      "post_date": "04/22/2020 14:05:28",
      "content": "<p>Try these out.</p>\n\n<p><a href=\"https://www.tensorflow.org/api_docs/python/tf/data/experimental/choose_from_datasets\">https://www.tensorflow.org/api_docs/python/tf/data/experimental/choose_from_datasets</a></p>\n\n<p>or</p>\n\n<p><a href=\"https://www.tensorflow.org/api_docs/python/tf/data/experimental/sample_from_datasets\">https://www.tensorflow.org/api_docs/python/tf/data/experimental/sample_from_datasets</a></p>\n\n<p>or just tf.data.Dataset.zip</p>\n\n<p>I put together some quick working examples of each of these in <a href=\"https://www.kaggle.com/calebeverett/combining-dataset-examples\">combining-dataset-examples</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 817059,
          "author_name": "rashanarshad",
          "author_url": "",
          "post_date": "04/22/2020 19:55:41",
          "content": "<p>thank you for the links Ive been looking through them right now, but I have a up couple follow up questions if you don't mind.</p>\n\n<p>So this is the algorithm im using to generate the pairs, this is a barely tweaked version from some tutorial \n```\ndef make_pairs(x,y):\n    num_classes = max(y) + 1\n    digit_indices = [np.where(y == i)[0] for i in range(num_classes)]</p>\n\n<h1>pairs = []</h1>\n\n<pre><code>left_images = []\nright_images = []\nlabels = []\n\nfor idx1 in range(len(x)):\n    #add matching example\n    x1 = x[idx1]\n    label1 = y[idx1]\n    idx2 = random.choice(digit_indices[label1])\n    x2 = x[idx2]\n</code></pre>\n\n<h1>pairs+= [[x1, x2]]</h1>\n\n<pre><code>    left_images.append(x1)\n    right_images.append(x2)\n    labels += [1]\n\n    # add a not matching example\n    label2 = random.randint(0, num_classes-1)\n    while label2 == label1:\n        label2 = random.randint(0, num_classes-1)\n\n    idx2 = random.choice(digit_indices[label2])\n    x2 = x[idx2]\n</code></pre>\n\n<h1>pairs += [[x1,x2]]</h1>\n\n<pre><code>    left_images.append(x1)\n    right_images.append(x2)\n    labels += [0]\n\nreturn left_images, right_images, labels\n</code></pre>\n\n<p><code>\n</code>\ndef get_train(image_size):\n    images, labels = make_single_dataset(image_size, TRAINING_FILENAMES,labeled=True)\n    left_imgs, right_imgs, labels = make_pairs(images, labels)    </p>\n\n<pre><code>del images\ntrain_dict = {'input_1':left_imgs,\n              'input_2': right_imgs}\n\nout_dict = {'output_1':labels}\n\ndel left_imgs, right_imgs, labels\nprint('dictionaries generated')\ntrain_dict =  tf.data.Dataset.from_tensor_slices(train_dict)\nout_dict =  tf.data.Dataset.from_tensor_slices(out_dict)\n\nprint('tensors generated')\n\ntrain_dict = tf.data.Dataset.zip((train_dict, out_dict))\n\ndel out_dict\nprint('tensors combined')\ntrain_dict = train_dict.repeat()\ntrain_dict = train_dict.batch(BATCH_SIZE)\ntrain_dict = train_dict.prefetch(AUTO) # prefetch next batch while training (autotune prefetch buffer size)\n\n\n\n\n\nreturn train_dict\n</code></pre>\n\n<p>```</p>\n\n<p>So I generate a list of all images and labels by iterating over a dataset that the files were loaded into.  I pass those to a make_pairs function that goes through every image and gets another image of the same class and an image of the different class  class.  these are stored in three lists (left images, right images and labels). I store these into dictionaries with the corresponding label for the model then load them into tensors.    Later I zip these two tensors using data.zip</p>\n\n<p>Do you think I can leverage the functions you mentioned to make this process more efficient ?  By the time I'm finished loading the training set's input and output dictionaries I'm using almost 10gbs of ram.  I'm deleting the unused objects (correctly, I hope) but that only helped so much </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 817069,
          "author_name": "rashanarshad",
          "author_url": "",
          "post_date": "04/22/2020 20:06:57",
          "content": "<p>another quick thought I had is if there simply isn't enough memory to do this with the whole dataset, I can possibly use sample by dataset for n amount of images where n &lt; Len(training_files). by calculating proper class weights I may even be able to generate a balanced , albeit , smaller, dataset. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "816319": "So essentially the dataset will be twice the size of the original training set (same length, two columns)\n\nI have an algorithm to generate the pairs how I want them (in the sense of distribution of classes) but I am having trouble fitting it all in memory.\n\nCurrently I am loading in the data as a dataset according to the starter notebook.\n\nBut to generate the pairs I am iterating over this dataset to split it into a list of images and a list of the corresponding labels.\n\nthen I am using these two generated lists to generate a list of pairs of images and list of labels signifiying if the pairs are of the same class or not.  \n\nTHEN I am storing them in a dictionary and converting them to tensors \n\nAt the end of this process I am already using 6gbs of ram just for the training set.",
    "816679": "Try these out.\n\nhttps://www.tensorflow.org/api_docs/python/tf/data/experimental/choose_from_datasets\n\nor\n\nhttps://www.tensorflow.org/api_docs/python/tf/data/experimental/sample_from_datasets\n\nor just tf.data.Dataset.zip\n\nI put together some quick working examples of each of these in [combining-dataset-examples](https://www.kaggle.com/calebeverett/combining-dataset-examples).",
    "817059": "thank you for the links Ive been looking through them right now, but I have a up couple follow up questions if you don't mind.\n\nSo this is the algorithm im using to generate the pairs, this is a barely tweaked version from some tutorial \n```\ndef make_pairs(x,y):\n    num_classes = max(y) + 1\n    digit_indices = [np.where(y == i)[0] for i in range(num_classes)]\n#     pairs = []\n    left_images = []\n    right_images = []\n    labels = []\n    \n    for idx1 in range(len(x)):\n        #add matching example\n        x1 = x[idx1]\n        label1 = y[idx1]\n        idx2 = random.choice(digit_indices[label1])\n        x2 = x[idx2]\n        \n#         pairs+= [[x1, x2]]\n        left_images.append(x1)\n        right_images.append(x2)\n        labels += [1]\n        \n        # add a not matching example\n        label2 = random.randint(0, num_classes-1)\n        while label2 == label1:\n            label2 = random.randint(0, num_classes-1)\n            \n        idx2 = random.choice(digit_indices[label2])\n        x2 = x[idx2]\n        \n#         pairs += [[x1,x2]]\n        left_images.append(x1)\n        right_images.append(x2)\n        labels += [0]\n        \n    return left_images, right_images, labels\n```\n```\ndef get_train(image_size):\n    images, labels = make_single_dataset(image_size, TRAINING_FILENAMES,labeled=True)\n    left_imgs, right_imgs, labels = make_pairs(images, labels)    \n    \n    del images\n    train_dict = {'input_1':left_imgs,\n                  'input_2': right_imgs}\n\n    out_dict = {'output_1':labels}\n    \n    del left_imgs, right_imgs, labels\n    print('dictionaries generated')\n    train_dict =  tf.data.Dataset.from_tensor_slices(train_dict)\n    out_dict =  tf.data.Dataset.from_tensor_slices(out_dict)\n\n    print('tensors generated')\n\n    train_dict = tf.data.Dataset.zip((train_dict, out_dict))\n    \n    del out_dict\n    print('tensors combined')\n    train_dict = train_dict.repeat()\n    train_dict = train_dict.batch(BATCH_SIZE)\n    train_dict = train_dict.prefetch(AUTO) # prefetch next batch while training (autotune prefetch buffer size)\n\n\n\n\n  \n    return train_dict\n```\n\n\nSo I generate a list of all images and labels by iterating over a dataset that the files were loaded into.  I pass those to a make_pairs function that goes through every image and gets another image of the same class and an image of the different class  class.  these are stored in three lists (left images, right images and labels). I store these into dictionaries with the corresponding label for the model then load them into tensors.    Later I zip these two tensors using data.zip\n\nDo you think I can leverage the functions you mentioned to make this process more efficient ?  By the time I'm finished loading the training set's input and output dictionaries I'm using almost 10gbs of ram.  I'm deleting the unused objects (correctly, I hope) but that only helped so much",
    "817069": "another quick thought I had is if there simply isn't enough memory to do this with the whole dataset, I can possibly use sample by dataset for n amount of images where n &lt; Len(training_files). by calculating proper class weights I may even be able to generate a balanced , albeit , smaller, dataset."
  },
  "source": "meta"
}