{
  "id": 142953,
  "title": "Does anyone have any experience with Siamese networks? I'm having trouble structuring the data",
  "url": "/competitions/flower-classification-with-tpus/discussion/142953",
  "author_name": "",
  "post_date": "2020-04-13T03:14:37.776903700Z",
  "votes": 1,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Additionally structuring it to pass through to the model fitting</p>\n\n<p>I can follow the getting started notebook to load the data into a dataset (which is a filename mapped to a decoding function which uses the the filename to load the image and it's respective label) </p>\n\n<p>but I don't really know how to structure the data set so it is image_1, image_2, label (being a binary label  if they are of the same class or not)</p>\n\n<p>I can build a model that accepts two inputs and out puts the binary classification and I understand the logic of how to construct the test set to generate the necessary predictions, but struggling with actually putting it to code</p>\n\n<p>So simply, is there someone who is familiar enough with tf datasets who can guide me toward structuring my dataset in the format I described? </p>\n\n<p>Thank You!</p>",
  "messages": [
    {
      "id": "805750",
      "postDate": "04/13/2020 03:14:37",
      "content": "<p>Additionally structuring it to pass through to the model fitting</p>\n\n<p>I can follow the getting started notebook to load the data into a dataset (which is a filename mapped to a decoding function which uses the the filename to load the image and it's respective label) </p>\n\n<p>but I don't really know how to structure the data set so it is image_1, image_2, label (being a binary label  if they are of the same class or not)</p>\n\n<p>I can build a model that accepts two inputs and out puts the binary classification and I understand the logic of how to construct the test set to generate the necessary predictions, but struggling with actually putting it to code</p>\n\n<p>So simply, is there someone who is familiar enough with tf datasets who can guide me toward structuring my dataset in the format I described? </p>\n\n<p>Thank You!</p>",
      "rawMarkdown": "Additionally structuring it to pass through to the model fitting\n\nI can follow the getting started notebook to load the data into a dataset (which is a filename mapped to a decoding function which uses the the filename to load the image and it's respective label) \n\nbut I don't really know how to structure the data set so it is image_1, image_2, label (being a binary label  if they are of the same class or not)\n\nI can build a model that accepts two inputs and out puts the binary classification and I understand the logic of how to construct the test set to generate the necessary predictions, but struggling with actually putting it to code\n\n\nSo simply, is there someone who is familiar enough with tf datasets who can guide me toward structuring my dataset in the format I described? \n\nThank You!",
      "votes": null
    },
    {
      "id": "807934",
      "postDate": "04/15/2020 03:52:02",
      "content": "<p>Use two dictionaries as a tuple, one for train data and one for labels. Here is an example from an NLP problem. The variables <code>ids</code>, <code>att</code>, <code>tok</code>, <code>tar1</code>, and <code>tar2</code> are numpy arrays. You'll need to update this to your situation.</p>\n\n<pre><code>    tf.data.Dataset\n        .from_tensor_slices( ({'input1':ids,'input2':att,'input3':tok}, {'output1':tar1,'output2':tar2}) )\n        .repeat()\n        .shuffle(2048)\n        .batch(BATCH_SIZE)\n        .prefetch(AUTO) \n</code></pre>",
      "rawMarkdown": "Use two dictionaries as a tuple, one for train data and one for labels. Here is an example from an NLP problem. The variables `ids`, `att`, `tok`, `tar1`, and `tar2` are numpy arrays. You'll need to update this to your situation.\n\n        \n        tf.data.Dataset\n            .from_tensor_slices( ({'input1':ids,'input2':att,'input3':tok}, {'output1':tar1,'output2':tar2}) )\n            .repeat()\n            .shuffle(2048)\n            .batch(BATCH_SIZE)\n            .prefetch(AUTO)",
      "votes": null
    },
    {
      "id": "809127",
      "postDate": "04/15/2020 22:02:51",
      "content": "<p>Hey Chris</p>\n\n<p>Appreciate the response, this is very helpful! </p>\n\n<p>Quick question for clarification: the<code>ids</code>, <code>att</code>, and <code>Tok</code> are three inputs, would this correspond to <code>image_set_1</code>, <code>image_set_2</code>, <code>similarity_label</code>? additionally, what are the outputs in this instance and would I need outputs in my implementation? This is within the dataset after all.  The model would only have one output as well to my understanding. </p>\n\n<p>or would it be something like this: first dictionary with two elements, one for each set of images. second dictionary with one element, the similarity labels.  Unless I am understanding siamese networks wrong altogether and I still need to send in the original labels for each of the images? </p>\n\n<p>edit 1:\nI gave it a quick attempt to implement it with assumptions from my questions above and Im getting an error: \n<code>ValueError: No data provided for \"input_1\". Need data for each key in: ['input_1', 'input_2']</code></p>\n\n<p>Im assuming my assumptions were wrong.  I structure it as two dictionaries, first dictionary with an array of the images extracted like this:\n<code>dataset = load_dataset(VALIDATION_FILENAMES, labeled=True, ordered=ordered)</code>\n <code>for image, label in dataset:</code>\n  <code>set_1.append(image)</code>\n   <code>l_1.append(label)</code></p>\n\n<p>did this for two datasets after they were shuffled and then constructed a third label array based on if the two labels are the same or not from the corresponding shuffled images.  </p>\n\n<p>edit2:</p>\n\n<p>Now I am getting this error after structuring dictionaries as:\n<code>data_dict = {'input_1' : set_1, 'input_2': set_2 , 'p_labels':pair_labels}\n    label_dict = {'l_1': l_1, 'l_2':l_2}</code></p>\n\n<p>and am getting this error:\n<code>ValueError: No data provided for \"dense_1\". Need data for each key in: ['dense_1']</code></p>\n\n<p>Seems like Im labeling the dictionary wrong but I also think I am not passing in the data correctly either.  </p>\n\n<p>edit3:\nIm getting it to train! however I dont think my dictionary and the way I am passing through my data is making any sense. </p>\n\n<p>Thank you for any help in advance!</p>",
      "rawMarkdown": "Hey Chris\n\nAppreciate the response, this is very helpful! \n\nQuick question for clarification: the` ids`, `att`, and `Tok` are three inputs, would this correspond to `image_set_1`, `image_set_2`, `similarity_label `? additionally, what are the outputs in this instance and would I need outputs in my implementation? This is within the dataset after all.  The model would only have one output as well to my understanding. \n\nor would it be something like this: first dictionary with two elements, one for each set of images. second dictionary with one element, the similarity labels.  Unless I am understanding siamese networks wrong altogether and I still need to send in the original labels for each of the images? \n\nedit 1:\nI gave it a quick attempt to implement it with assumptions from my questions above and Im getting an error: \n`ValueError: No data provided for \"input_1\". Need data for each key in: ['input_1', 'input_2']`\n\nIm assuming my assumptions were wrong.  I structure it as two dictionaries, first dictionary with an array of the images extracted like this:\n`dataset = load_dataset(VALIDATION_FILENAMES, labeled=True, ordered=ordered)`\n `   for image, label in dataset:`\n  `      set_1.append(image)`\n   `     l_1.append(label)`\n\ndid this for two datasets after they were shuffled and then constructed a third label array based on if the two labels are the same or not from the corresponding shuffled images.  \n\nedit2:\n\nNow I am getting this error after structuring dictionaries as:\n`data_dict = {'input_1' : set_1, 'input_2': set_2 , 'p_labels':pair_labels}\n    label_dict = {'l_1': l_1, 'l_2':l_2}`\n\nand am getting this error:\n`    ValueError: No data provided for \"dense_1\". Need data for each key in: ['dense_1']`\n\nSeems like Im labeling the dictionary wrong but I also think I am not passing in the data correctly either.  \n\nedit3:\nIm getting it to train! however I dont think my dictionary and the way I am passing through my data is making any sense. \n\nThank you for any help in advance!",
      "votes": null
    },
    {
      "id": "809169",
      "postDate": "04/15/2020 23:20:19",
      "content": "<p>My other comment was getting too long but to simplify I have two follow up questions:</p>\n\n<ol>\n<li>Still a bit confused on the ordering / structuring of the data being passed through to train, this is what I currently have (that passes through to the model and begins training):\n<code>data_dict = {'input_1' : set_1, 'input_2': set_2}</code>\n<code>label_dict = {'dense_1' : pair_labels}</code></li>\n<li>Any guidance on how to use the model to make predictions? Am I creating a dataset with an image from the unseen set paired with an example of each of the species then take the most similar value? </li>\n</ol>\n\n<p>Thank You!</p>",
      "rawMarkdown": "My other comment was getting too long but to simplify I have two follow up questions:\n\n1. Still a bit confused on the ordering / structuring of the data being passed through to train, this is what I currently have (that passes through to the model and begins training):\n`    data_dict = {'input_1' : set_1, 'input_2': set_2}`\n`    label_dict = {'dense_1' : pair_labels}`\n2. Any guidance on how to use the model to make predictions? Am I creating a dataset with an image from the unseen set paired with an example of each of the species then take the most similar value? \n\nThank You!",
      "votes": null
    },
    {
      "id": "809172",
      "postDate": "04/15/2020 23:26:09",
      "content": "<p>My example was 3 inputs and 2 outputs. If your model has 2 inputs and 1 output, then the dataset would be </p>\n\n<pre><code>    tf.data.Dataset.from_tensor_slices( ({'input1':images1,'input2':images2}, {'output1':targets}) )\n</code></pre>\n\n<p>One dictionary with 2 inputs, and one dictionary with 1 output. (or just the output without second dictionary). Here <code>images1</code>, <code>images2</code>, and <code>targets</code> are all numpy arrays. </p>",
      "rawMarkdown": "My example was 3 inputs and 2 outputs. If your model has 2 inputs and 1 output, then the dataset would be \n\n        tf.data.Dataset.from_tensor_slices( ({'input1':images1,'input2':images2}, {'output1':targets}) )\n\nOne dictionary with 2 inputs, and one dictionary with 1 output. (or just the output without second dictionary). Here `images1`, `images2`, and `targets` are all numpy arrays.",
      "votes": null
    },
    {
      "id": "809174",
      "postDate": "04/15/2020 23:29:20",
      "content": "<blockquote>\n  <p>ValueError: No data provided for \"<code>dense_1</code>\". Need data for each key in: ['<code>dense_1</code>']</p>\n</blockquote>\n\n<p>In your model, you need to name the output layers to match the names in the dictionary. So for example if your last layer is a dense layer like so:</p>\n\n<pre><code>x = tf.keras.layers.Dense(1, name='output1')(x)\n</code></pre>\n\n<p>If you don't provide a name, it is called <code>dense_1</code>.</p>",
      "rawMarkdown": "&gt; ValueError: No data provided for \"`dense_1`\". Need data for each key in: ['`dense_1`']\n\nIn your model, you need to name the output layers to match the names in the dictionary. So for example if your last layer is a dense layer like so:\n\n    x = tf.keras.layers.Dense(1, name='output1')(x)\n\nIf you don't provide a name, it is called `dense_1`.",
      "votes": null
    },
    {
      "id": "809176",
      "postDate": "04/15/2020 23:30:54",
      "content": "<p>And name the inputs in your model</p>\n\n<pre><code>inp1 = tf.keras.layers.Input(DIM, name='input1')\ninp2 = tf.keras.layers.Input(DIM, name='input2')\n</code></pre>\n\n<p>...</p>\n\n<pre><code>model = tf.keras.models.Model(inputs=[inp1,inp2], outputs=[x])\n</code></pre>",
      "rawMarkdown": "And name the inputs in your model\n\n    inp1 = tf.keras.layers.Input(DIM, name='input1')\n    inp2 = tf.keras.layers.Input(DIM, name='input2')\n\n...\n\n    model = tf.keras.models.Model(inputs=[inp1,inp2], outputs=[x])",
      "votes": null
    },
    {
      "id": "809183",
      "postDate": "04/15/2020 23:43:03",
      "content": "<p>Awesome, this all helped greatly! am able to train my model now and can start simplifying the pipeline.  </p>\n\n<p>For predicting the class on the unseen dataset, is my intuition correct? Am I creating a dataset with an image from the unseen set paired with an example from each of the species then take the most similar value?</p>",
      "rawMarkdown": "Awesome, this all helped greatly! am able to train my model now and can start simplifying the pipeline.  \n\nFor predicting the class on the unseen dataset, is my intuition correct? Am I creating a dataset with an image from the unseen set paired with an example from each of the species then take the most similar value?",
      "votes": null
    },
    {
      "id": "809191",
      "postDate": "04/15/2020 23:57:43",
      "content": "<p>If your goal is to find similar images, there's another way besides Siamese network. You just feed all the train images and all the test images into a pretrained ImageNet CNN (for example EfficientNet). Then you apply <code>GlobalAveragePooling2D</code> and output the result. (Don't apply Dense layers). Afterward you have a vector of size 2000 (approximately) for each flower. Then apply kNN (using RAPIDS' GPU implementation) on those flower vectors. That's how you do similarity image retrieval. </p>\n\n<p>(I don't have experience with Siamese image retrieval but I think you would want to compare each test image to more than 1 training image example).</p>",
      "rawMarkdown": "If your goal is to find similar images, there's another way besides Siamese network. You just feed all the train images and all the test images into a pretrained ImageNet CNN (for example EfficientNet). Then you apply `GlobalAveragePooling2D` and output the result. (Don't apply Dense layers). Afterward you have a vector of size 2000 (approximately) for each flower. Then apply kNN (using RAPIDS' GPU implementation) on those flower vectors. That's how you do similarity image retrieval. \n\n(I don't have experience with Siamese image retrieval but I think you would want to compare each test image to more than 1 training image example).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 807934,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "04/15/2020 03:52:02",
      "content": "<p>Use two dictionaries as a tuple, one for train data and one for labels. Here is an example from an NLP problem. The variables <code>ids</code>, <code>att</code>, <code>tok</code>, <code>tar1</code>, and <code>tar2</code> are numpy arrays. You'll need to update this to your situation.</p>\n\n<pre><code>    tf.data.Dataset\n        .from_tensor_slices( ({'input1':ids,'input2':att,'input3':tok}, {'output1':tar1,'output2':tar2}) )\n        .repeat()\n        .shuffle(2048)\n        .batch(BATCH_SIZE)\n        .prefetch(AUTO) \n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 809127,
          "author_name": "rashanarshad",
          "author_url": "",
          "post_date": "04/15/2020 22:02:51",
          "content": "<p>Hey Chris</p>\n\n<p>Appreciate the response, this is very helpful! </p>\n\n<p>Quick question for clarification: the<code>ids</code>, <code>att</code>, and <code>Tok</code> are three inputs, would this correspond to <code>image_set_1</code>, <code>image_set_2</code>, <code>similarity_label</code>? additionally, what are the outputs in this instance and would I need outputs in my implementation? This is within the dataset after all.  The model would only have one output as well to my understanding. </p>\n\n<p>or would it be something like this: first dictionary with two elements, one for each set of images. second dictionary with one element, the similarity labels.  Unless I am understanding siamese networks wrong altogether and I still need to send in the original labels for each of the images? </p>\n\n<p>edit 1:\nI gave it a quick attempt to implement it with assumptions from my questions above and Im getting an error: \n<code>ValueError: No data provided for \"input_1\". Need data for each key in: ['input_1', 'input_2']</code></p>\n\n<p>Im assuming my assumptions were wrong.  I structure it as two dictionaries, first dictionary with an array of the images extracted like this:\n<code>dataset = load_dataset(VALIDATION_FILENAMES, labeled=True, ordered=ordered)</code>\n <code>for image, label in dataset:</code>\n  <code>set_1.append(image)</code>\n   <code>l_1.append(label)</code></p>\n\n<p>did this for two datasets after they were shuffled and then constructed a third label array based on if the two labels are the same or not from the corresponding shuffled images.  </p>\n\n<p>edit2:</p>\n\n<p>Now I am getting this error after structuring dictionaries as:\n<code>data_dict = {'input_1' : set_1, 'input_2': set_2 , 'p_labels':pair_labels}\n    label_dict = {'l_1': l_1, 'l_2':l_2}</code></p>\n\n<p>and am getting this error:\n<code>ValueError: No data provided for \"dense_1\". Need data for each key in: ['dense_1']</code></p>\n\n<p>Seems like Im labeling the dictionary wrong but I also think I am not passing in the data correctly either.  </p>\n\n<p>edit3:\nIm getting it to train! however I dont think my dictionary and the way I am passing through my data is making any sense. </p>\n\n<p>Thank you for any help in advance!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 809169,
          "author_name": "rashanarshad",
          "author_url": "",
          "post_date": "04/15/2020 23:20:19",
          "content": "<p>My other comment was getting too long but to simplify I have two follow up questions:</p>\n\n<ol>\n<li>Still a bit confused on the ordering / structuring of the data being passed through to train, this is what I currently have (that passes through to the model and begins training):\n<code>data_dict = {'input_1' : set_1, 'input_2': set_2}</code>\n<code>label_dict = {'dense_1' : pair_labels}</code></li>\n<li>Any guidance on how to use the model to make predictions? Am I creating a dataset with an image from the unseen set paired with an example of each of the species then take the most similar value? </li>\n</ol>\n\n<p>Thank You!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 809172,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "04/15/2020 23:26:09",
          "content": "<p>My example was 3 inputs and 2 outputs. If your model has 2 inputs and 1 output, then the dataset would be </p>\n\n<pre><code>    tf.data.Dataset.from_tensor_slices( ({'input1':images1,'input2':images2}, {'output1':targets}) )\n</code></pre>\n\n<p>One dictionary with 2 inputs, and one dictionary with 1 output. (or just the output without second dictionary). Here <code>images1</code>, <code>images2</code>, and <code>targets</code> are all numpy arrays. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 809174,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "04/15/2020 23:29:20",
          "content": "<blockquote>\n  <p>ValueError: No data provided for \"<code>dense_1</code>\". Need data for each key in: ['<code>dense_1</code>']</p>\n</blockquote>\n\n<p>In your model, you need to name the output layers to match the names in the dictionary. So for example if your last layer is a dense layer like so:</p>\n\n<pre><code>x = tf.keras.layers.Dense(1, name='output1')(x)\n</code></pre>\n\n<p>If you don't provide a name, it is called <code>dense_1</code>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 809176,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "04/15/2020 23:30:54",
          "content": "<p>And name the inputs in your model</p>\n\n<pre><code>inp1 = tf.keras.layers.Input(DIM, name='input1')\ninp2 = tf.keras.layers.Input(DIM, name='input2')\n</code></pre>\n\n<p>...</p>\n\n<pre><code>model = tf.keras.models.Model(inputs=[inp1,inp2], outputs=[x])\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 809183,
          "author_name": "rashanarshad",
          "author_url": "",
          "post_date": "04/15/2020 23:43:03",
          "content": "<p>Awesome, this all helped greatly! am able to train my model now and can start simplifying the pipeline.  </p>\n\n<p>For predicting the class on the unseen dataset, is my intuition correct? Am I creating a dataset with an image from the unseen set paired with an example from each of the species then take the most similar value?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 809191,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "04/15/2020 23:57:43",
          "content": "<p>If your goal is to find similar images, there's another way besides Siamese network. You just feed all the train images and all the test images into a pretrained ImageNet CNN (for example EfficientNet). Then you apply <code>GlobalAveragePooling2D</code> and output the result. (Don't apply Dense layers). Afterward you have a vector of size 2000 (approximately) for each flower. Then apply kNN (using RAPIDS' GPU implementation) on those flower vectors. That's how you do similarity image retrieval. </p>\n\n<p>(I don't have experience with Siamese image retrieval but I think you would want to compare each test image to more than 1 training image example).</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "805750": "Additionally structuring it to pass through to the model fitting\n\nI can follow the getting started notebook to load the data into a dataset (which is a filename mapped to a decoding function which uses the the filename to load the image and it's respective label) \n\nbut I don't really know how to structure the data set so it is image_1, image_2, label (being a binary label  if they are of the same class or not)\n\nI can build a model that accepts two inputs and out puts the binary classification and I understand the logic of how to construct the test set to generate the necessary predictions, but struggling with actually putting it to code\n\n\nSo simply, is there someone who is familiar enough with tf datasets who can guide me toward structuring my dataset in the format I described? \n\nThank You!",
    "807934": "Use two dictionaries as a tuple, one for train data and one for labels. Here is an example from an NLP problem. The variables `ids`, `att`, `tok`, `tar1`, and `tar2` are numpy arrays. You'll need to update this to your situation.\n\n        \n        tf.data.Dataset\n            .from_tensor_slices( ({'input1':ids,'input2':att,'input3':tok}, {'output1':tar1,'output2':tar2}) )\n            .repeat()\n            .shuffle(2048)\n            .batch(BATCH_SIZE)\n            .prefetch(AUTO)",
    "809127": "Hey Chris\n\nAppreciate the response, this is very helpful! \n\nQuick question for clarification: the` ids`, `att`, and `Tok` are three inputs, would this correspond to `image_set_1`, `image_set_2`, `similarity_label `? additionally, what are the outputs in this instance and would I need outputs in my implementation? This is within the dataset after all.  The model would only have one output as well to my understanding. \n\nor would it be something like this: first dictionary with two elements, one for each set of images. second dictionary with one element, the similarity labels.  Unless I am understanding siamese networks wrong altogether and I still need to send in the original labels for each of the images? \n\nedit 1:\nI gave it a quick attempt to implement it with assumptions from my questions above and Im getting an error: \n`ValueError: No data provided for \"input_1\". Need data for each key in: ['input_1', 'input_2']`\n\nIm assuming my assumptions were wrong.  I structure it as two dictionaries, first dictionary with an array of the images extracted like this:\n`dataset = load_dataset(VALIDATION_FILENAMES, labeled=True, ordered=ordered)`\n `   for image, label in dataset:`\n  `      set_1.append(image)`\n   `     l_1.append(label)`\n\ndid this for two datasets after they were shuffled and then constructed a third label array based on if the two labels are the same or not from the corresponding shuffled images.  \n\nedit2:\n\nNow I am getting this error after structuring dictionaries as:\n`data_dict = {'input_1' : set_1, 'input_2': set_2 , 'p_labels':pair_labels}\n    label_dict = {'l_1': l_1, 'l_2':l_2}`\n\nand am getting this error:\n`    ValueError: No data provided for \"dense_1\". Need data for each key in: ['dense_1']`\n\nSeems like Im labeling the dictionary wrong but I also think I am not passing in the data correctly either.  \n\nedit3:\nIm getting it to train! however I dont think my dictionary and the way I am passing through my data is making any sense. \n\nThank you for any help in advance!",
    "809169": "My other comment was getting too long but to simplify I have two follow up questions:\n\n1. Still a bit confused on the ordering / structuring of the data being passed through to train, this is what I currently have (that passes through to the model and begins training):\n`    data_dict = {'input_1' : set_1, 'input_2': set_2}`\n`    label_dict = {'dense_1' : pair_labels}`\n2. Any guidance on how to use the model to make predictions? Am I creating a dataset with an image from the unseen set paired with an example of each of the species then take the most similar value? \n\nThank You!",
    "809172": "My example was 3 inputs and 2 outputs. If your model has 2 inputs and 1 output, then the dataset would be \n\n        tf.data.Dataset.from_tensor_slices( ({'input1':images1,'input2':images2}, {'output1':targets}) )\n\nOne dictionary with 2 inputs, and one dictionary with 1 output. (or just the output without second dictionary). Here `images1`, `images2`, and `targets` are all numpy arrays.",
    "809174": "&gt; ValueError: No data provided for \"`dense_1`\". Need data for each key in: ['`dense_1`']\n\nIn your model, you need to name the output layers to match the names in the dictionary. So for example if your last layer is a dense layer like so:\n\n    x = tf.keras.layers.Dense(1, name='output1')(x)\n\nIf you don't provide a name, it is called `dense_1`.",
    "809176": "And name the inputs in your model\n\n    inp1 = tf.keras.layers.Input(DIM, name='input1')\n    inp2 = tf.keras.layers.Input(DIM, name='input2')\n\n...\n\n    model = tf.keras.models.Model(inputs=[inp1,inp2], outputs=[x])",
    "809183": "Awesome, this all helped greatly! am able to train my model now and can start simplifying the pipeline.  \n\nFor predicting the class on the unseen dataset, is my intuition correct? Am I creating a dataset with an image from the unseen set paired with an example from each of the species then take the most similar value?",
    "809191": "If your goal is to find similar images, there's another way besides Siamese network. You just feed all the train images and all the test images into a pretrained ImageNet CNN (for example EfficientNet). Then you apply `GlobalAveragePooling2D` and output the result. (Don't apply Dense layers). Afterward you have a vector of size 2000 (approximately) for each flower. Then apply kNN (using RAPIDS' GPU implementation) on those flower vectors. That's how you do similarity image retrieval. \n\n(I don't have experience with Siamese image retrieval but I think you would want to compare each test image to more than 1 training image example)."
  },
  "source": "meta"
}