{
  "id": 318260,
  "title": "A truly spectacularly failed approach",
  "url": "/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/discussion/318260",
  "author_name": "",
  "post_date": "2022-04-11T11:28:24.074252800Z",
  "votes": 7,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>I had a possibly interesting path that I went down with this that completely failed. Not just a little failure - a 0.000 on validation failure! It's probably obvious to those with more experience why it failed, but I thought I'd share it regardless.</p>\n<p>So here was the idea: How do we actually compare two hotel images and decide that it's the same hotel? I suspect a lot of it will come down to common furnishings, building features such as the architraves, carpeting, etc. Hotels often standardise these things across rooms, to provide a common experience for their guests.</p>\n<p>So, how do we find these common features in rooms? Obviously a neural network can be trained to learn associations of feature to hotel, but do we even need to do that? Could we just find bits of the image that contain the same thing?</p>\n<p>I decided to test this out by indexing and searching vectors extracted from a pre-trained network - in this case EfficientNet-B0. The idea behind this was that the later layers of the network should contain semantic information, right? Maybe not quite as clear as lamp, books, cushion, etc, but something about the shapes and textures involved.</p>\n<p>By extracting those features for every image in the training set (minus a validation group), indexing them, and then searching for related vectors to those in a target image, can I simply count the hotels the nearest neighbours came from as \"votes\" for a winner?</p>\n<p>Well, as it turns out, no. I'd be very interested in other people's takes on why that might be the case, but I suspect the features are just not a dense vector of semantic information. There may just be too much noise left in that layer, putting every vector at 90 degrees to every other.</p>\n<p>I've tidied up my notebooks, to show the path taken to the spectacular zero result, and share the intermediate results for anyone who wants to see (best read in order):</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/prubyg/hotel-id-vector-extraction\" target=\"_blank\">https://www.kaggle.com/prubyg/hotel-id-vector-extraction</a></li>\n<li><a href=\"https://www.kaggle.com/prubyg/hotel-id-vector-indexing\" target=\"_blank\">https://www.kaggle.com/prubyg/hotel-id-vector-indexing</a></li>\n<li><a href=\"https://www.kaggle.com/prubyg/hotel-id-vector-validation\" target=\"_blank\">https://www.kaggle.com/prubyg/hotel-id-vector-validation</a></li>\n</ul>",
  "messages": [
    {
      "id": "1752073",
      "postDate": "04/11/2022 11:28:24",
      "content": "<p>Hi all,</p>\n<p>I had a possibly interesting path that I went down with this that completely failed. Not just a little failure - a 0.000 on validation failure! It's probably obvious to those with more experience why it failed, but I thought I'd share it regardless.</p>\n<p>So here was the idea: How do we actually compare two hotel images and decide that it's the same hotel? I suspect a lot of it will come down to common furnishings, building features such as the architraves, carpeting, etc. Hotels often standardise these things across rooms, to provide a common experience for their guests.</p>\n<p>So, how do we find these common features in rooms? Obviously a neural network can be trained to learn associations of feature to hotel, but do we even need to do that? Could we just find bits of the image that contain the same thing?</p>\n<p>I decided to test this out by indexing and searching vectors extracted from a pre-trained network - in this case EfficientNet-B0. The idea behind this was that the later layers of the network should contain semantic information, right? Maybe not quite as clear as lamp, books, cushion, etc, but something about the shapes and textures involved.</p>\n<p>By extracting those features for every image in the training set (minus a validation group), indexing them, and then searching for related vectors to those in a target image, can I simply count the hotels the nearest neighbours came from as \"votes\" for a winner?</p>\n<p>Well, as it turns out, no. I'd be very interested in other people's takes on why that might be the case, but I suspect the features are just not a dense vector of semantic information. There may just be too much noise left in that layer, putting every vector at 90 degrees to every other.</p>\n<p>I've tidied up my notebooks, to show the path taken to the spectacular zero result, and share the intermediate results for anyone who wants to see (best read in order):</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/prubyg/hotel-id-vector-extraction\" target=\"_blank\">https://www.kaggle.com/prubyg/hotel-id-vector-extraction</a></li>\n<li><a href=\"https://www.kaggle.com/prubyg/hotel-id-vector-indexing\" target=\"_blank\">https://www.kaggle.com/prubyg/hotel-id-vector-indexing</a></li>\n<li><a href=\"https://www.kaggle.com/prubyg/hotel-id-vector-validation\" target=\"_blank\">https://www.kaggle.com/prubyg/hotel-id-vector-validation</a></li>\n</ul>",
      "rawMarkdown": "Hi all,\n\nI had a possibly interesting path that I went down with this that completely failed. Not just a little failure - a 0.000 on validation failure! It's probably obvious to those with more experience why it failed, but I thought I'd share it regardless.\n\nSo here was the idea: How do we actually compare two hotel images and decide that it's the same hotel? I suspect a lot of it will come down to common furnishings, building features such as the architraves, carpeting, etc. Hotels often standardise these things across rooms, to provide a common experience for their guests.\n\nSo, how do we find these common features in rooms? Obviously a neural network can be trained to learn associations of feature to hotel, but do we even need to do that? Could we just find bits of the image that contain the same thing?\n\nI decided to test this out by indexing and searching vectors extracted from a pre-trained network - in this case EfficientNet-B0. The idea behind this was that the later layers of the network should contain semantic information, right? Maybe not quite as clear as lamp, books, cushion, etc, but something about the shapes and textures involved.\n\nBy extracting those features for every image in the training set (minus a validation group), indexing them, and then searching for related vectors to those in a target image, can I simply count the hotels the nearest neighbours came from as \"votes\" for a winner?\n\nWell, as it turns out, no. I'd be very interested in other people's takes on why that might be the case, but I suspect the features are just not a dense vector of semantic information. There may just be too much noise left in that layer, putting every vector at 90 degrees to every other.\n\nI've tidied up my notebooks, to show the path taken to the spectacular zero result, and share the intermediate results for anyone who wants to see (best read in order):\n\n- https://www.kaggle.com/prubyg/hotel-id-vector-extraction\n- https://www.kaggle.com/prubyg/hotel-id-vector-indexing\n- https://www.kaggle.com/prubyg/hotel-id-vector-validation",
      "votes": null
    },
    {
      "id": "1752283",
      "postDate": "04/11/2022 15:30:17",
      "content": "<p>Interesting approach, the idea is valid and I believe it should work. It is similar to traditional image matching method just using raw CNN features instead of generating embeddings for similarity search (you can check my <a href=\"https://www.kaggle.com/code/michaln/hotel-id-starter-similarity-training/notebook\" target=\"_blank\">Hotel-ID starter - similarity</a> notebook).</p>\n<p>I didn't go through the notebooks in detail and I am not familiar with the Annoy library but I tried to replace <br>\n<code>val_df = pd.read_csv('/kaggle/input/hotel-id-vector-extraction/validation.csv')</code><br>\nwith <br>\n<code>val_df = pd.read_csv('/kaggle/input/hotel-id-vector-extraction/training.csv')</code><br>\nin your Hotel ID - Vector Validation notebook and it still get zero correct matches. I would expect to get them all correctly so there must be something wrong.</p>\n<p>I found one error in the Hotel ID - Vector Validation candidate finding part<br>\nYou should replace this:</p>\n<pre><code>for hotel, vs in enumerate(votes):\n    if vs &gt; max_votes:\n</code></pre>\n<p>with </p>\n<pre><code>for i, hotel in enumerate(votes.keys()):\n    vs = votes[hotel]\n    if vs &gt; max_votes:\n</code></pre>\n<p>because in your case the enumerate returns index into the <em>hotel</em> variable and hotel_id (key from the map) into the <em>vs</em> variable. Then it get at least some correct for the training data and few for valid dataset. But still it's not great so it would be worth checking why it can't classify even the training data correctly.</p>\n<p>I can imagine few ways you could improve the solution:</p>\n<ul>\n<li>You could train the model to classify the images first and then extract the features with the finetuned model. This should give the model a better chance to learn features important for the task.</li>\n<li>The orientation of the images might be important in this case so you might want to identify the ones with incorrect orientation and rotate them accordingly. </li>\n<li>And maybe there is a better way to compare the features. If i undestand it correctly you generate 320 features of size 8x8 for each image and then do knn search for each of the 8x8 points accross all features. Maybe it would be better to just unpack the 8x8 features into 64 values and do search on them or you could try to use something like PCA to reduce the dimension size.</li>\n</ul>",
      "rawMarkdown": "Interesting approach, the idea is valid and I believe it should work. It is similar to traditional image matching method just using raw CNN features instead of generating embeddings for similarity search (you can check my [Hotel-ID starter - similarity](https://www.kaggle.com/code/michaln/hotel-id-starter-similarity-training/notebook) notebook).\n\nI didn't go through the notebooks in detail and I am not familiar with the Annoy library but I tried to replace \n`val_df = pd.read_csv('/kaggle/input/hotel-id-vector-extraction/validation.csv')`\nwith \n`val_df = pd.read_csv('/kaggle/input/hotel-id-vector-extraction/training.csv')`\nin your Hotel ID - Vector Validation notebook and it still get zero correct matches. I would expect to get them all correctly so there must be something wrong.\n\nI found one error in the Hotel ID - Vector Validation candidate finding part\nYou should replace this:\n```\nfor hotel, vs in enumerate(votes):\n\tif vs > max_votes:\n```\n\nwith \n```\nfor i, hotel in enumerate(votes.keys()):\n\tvs = votes[hotel]\n\tif vs > max_votes:\n```\nbecause in your case the enumerate returns index into the *hotel* variable and hotel_id (key from the map) into the *vs* variable. Then it get at least some correct for the training data and few for valid dataset. But still it's not great so it would be worth checking why it can't classify even the training data correctly.\n\nI can imagine few ways you could improve the solution:\n- You could train the model to classify the images first and then extract the features with the finetuned model. This should give the model a better chance to learn features important for the task.\n- The orientation of the images might be important in this case so you might want to identify the ones with incorrect orientation and rotate them accordingly. \n- And maybe there is a better way to compare the features. If i undestand it correctly you generate 320 features of size 8x8 for each image and then do knn search for each of the 8x8 points accross all features. Maybe it would be better to just unpack the 8x8 features into 64 values and do search on them or you could try to use something like PCA to reduce the dimension size.",
      "votes": null
    },
    {
      "id": "1752559",
      "postDate": "04/11/2022 23:25:37",
      "content": "<p>Thank you, fixed that bug by changing it to votes.items(), which is what I meant to use. That improves it to get 10 of the validation set right.</p>\n<p>Good leads on training a classifier and fixing orientation, thank you.</p>\n<p>If I'm understanding correctly, your suggested method here is what I'm doing. I'm extracting 64 (8x8) 320-dimensional feature vectors from each image, which are what are being searched.</p>",
      "rawMarkdown": "Thank you, fixed that bug by changing it to votes.items(), which is what I meant to use. That improves it to get 10 of the validation set right.\n\nGood leads on training a classifier and fixing orientation, thank you.\n\nIf I'm understanding correctly, your suggested method here is what I'm doing. I'm extracting 64 (8x8) 320-dimensional feature vectors from each image, which are what are being searched.",
      "votes": null
    },
    {
      "id": "1752837",
      "postDate": "04/12/2022 07:22:25",
      "content": "<p>Well my suggestion was more like doing 320 searches on 64 value feature vector so you would be looking for images with similar values in each feature or just do one knn on all feature values at once. But maybe it would lead to the same result not sure (it's just searching in different dimension I guess). But if you do 320 searches for each feature you could try to identify which feature was important (lead to most correct predictions) and visualize it for fun or give them weights in final voting.</p>\n<p>Did you try some other method than Annoy? When I tried to classify training data, use only the base_transform (no occlusions) and K=1 it still didn't get very good results. Even though for KNN I would expect the nearest neighbor would be the input image itself.</p>",
      "rawMarkdown": "Well my suggestion was more like doing 320 searches on 64 value feature vector so you would be looking for images with similar values in each feature or just do one knn on all feature values at once. But maybe it would lead to the same result not sure (it's just searching in different dimension I guess). But if you do 320 searches for each feature you could try to identify which feature was important (lead to most correct predictions) and visualize it for fun or give them weights in final voting.\n\nDid you try some other method than Annoy? When I tried to classify training data, use only the base_transform (no occlusions) and K=1 it still didn't get very good results. Even though for KNN I would expect the nearest neighbor would be the input image itself.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1752283,
      "author_name": "michaln",
      "author_url": "",
      "post_date": "04/11/2022 15:30:17",
      "content": "<p>Interesting approach, the idea is valid and I believe it should work. It is similar to traditional image matching method just using raw CNN features instead of generating embeddings for similarity search (you can check my <a href=\"https://www.kaggle.com/code/michaln/hotel-id-starter-similarity-training/notebook\" target=\"_blank\">Hotel-ID starter - similarity</a> notebook).</p>\n<p>I didn't go through the notebooks in detail and I am not familiar with the Annoy library but I tried to replace <br>\n<code>val_df = pd.read_csv('/kaggle/input/hotel-id-vector-extraction/validation.csv')</code><br>\nwith <br>\n<code>val_df = pd.read_csv('/kaggle/input/hotel-id-vector-extraction/training.csv')</code><br>\nin your Hotel ID - Vector Validation notebook and it still get zero correct matches. I would expect to get them all correctly so there must be something wrong.</p>\n<p>I found one error in the Hotel ID - Vector Validation candidate finding part<br>\nYou should replace this:</p>\n<pre><code>for hotel, vs in enumerate(votes):\n    if vs &gt; max_votes:\n</code></pre>\n<p>with </p>\n<pre><code>for i, hotel in enumerate(votes.keys()):\n    vs = votes[hotel]\n    if vs &gt; max_votes:\n</code></pre>\n<p>because in your case the enumerate returns index into the <em>hotel</em> variable and hotel_id (key from the map) into the <em>vs</em> variable. Then it get at least some correct for the training data and few for valid dataset. But still it's not great so it would be worth checking why it can't classify even the training data correctly.</p>\n<p>I can imagine few ways you could improve the solution:</p>\n<ul>\n<li>You could train the model to classify the images first and then extract the features with the finetuned model. This should give the model a better chance to learn features important for the task.</li>\n<li>The orientation of the images might be important in this case so you might want to identify the ones with incorrect orientation and rotate them accordingly. </li>\n<li>And maybe there is a better way to compare the features. If i undestand it correctly you generate 320 features of size 8x8 for each image and then do knn search for each of the 8x8 points accross all features. Maybe it would be better to just unpack the 8x8 features into 64 values and do search on them or you could try to use something like PCA to reduce the dimension size.</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1752559,
          "author_name": "prubyg",
          "author_url": "",
          "post_date": "04/11/2022 23:25:37",
          "content": "<p>Thank you, fixed that bug by changing it to votes.items(), which is what I meant to use. That improves it to get 10 of the validation set right.</p>\n<p>Good leads on training a classifier and fixing orientation, thank you.</p>\n<p>If I'm understanding correctly, your suggested method here is what I'm doing. I'm extracting 64 (8x8) 320-dimensional feature vectors from each image, which are what are being searched.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1752837,
          "author_name": "michaln",
          "author_url": "",
          "post_date": "04/12/2022 07:22:25",
          "content": "<p>Well my suggestion was more like doing 320 searches on 64 value feature vector so you would be looking for images with similar values in each feature or just do one knn on all feature values at once. But maybe it would lead to the same result not sure (it's just searching in different dimension I guess). But if you do 320 searches for each feature you could try to identify which feature was important (lead to most correct predictions) and visualize it for fun or give them weights in final voting.</p>\n<p>Did you try some other method than Annoy? When I tried to classify training data, use only the base_transform (no occlusions) and K=1 it still didn't get very good results. Even though for KNN I would expect the nearest neighbor would be the input image itself.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1752073": "Hi all,\n\nI had a possibly interesting path that I went down with this that completely failed. Not just a little failure - a 0.000 on validation failure! It's probably obvious to those with more experience why it failed, but I thought I'd share it regardless.\n\nSo here was the idea: How do we actually compare two hotel images and decide that it's the same hotel? I suspect a lot of it will come down to common furnishings, building features such as the architraves, carpeting, etc. Hotels often standardise these things across rooms, to provide a common experience for their guests.\n\nSo, how do we find these common features in rooms? Obviously a neural network can be trained to learn associations of feature to hotel, but do we even need to do that? Could we just find bits of the image that contain the same thing?\n\nI decided to test this out by indexing and searching vectors extracted from a pre-trained network - in this case EfficientNet-B0. The idea behind this was that the later layers of the network should contain semantic information, right? Maybe not quite as clear as lamp, books, cushion, etc, but something about the shapes and textures involved.\n\nBy extracting those features for every image in the training set (minus a validation group), indexing them, and then searching for related vectors to those in a target image, can I simply count the hotels the nearest neighbours came from as \"votes\" for a winner?\n\nWell, as it turns out, no. I'd be very interested in other people's takes on why that might be the case, but I suspect the features are just not a dense vector of semantic information. There may just be too much noise left in that layer, putting every vector at 90 degrees to every other.\n\nI've tidied up my notebooks, to show the path taken to the spectacular zero result, and share the intermediate results for anyone who wants to see (best read in order):\n\n- https://www.kaggle.com/prubyg/hotel-id-vector-extraction\n- https://www.kaggle.com/prubyg/hotel-id-vector-indexing\n- https://www.kaggle.com/prubyg/hotel-id-vector-validation",
    "1752283": "Interesting approach, the idea is valid and I believe it should work. It is similar to traditional image matching method just using raw CNN features instead of generating embeddings for similarity search (you can check my [Hotel-ID starter - similarity](https://www.kaggle.com/code/michaln/hotel-id-starter-similarity-training/notebook) notebook).\n\nI didn't go through the notebooks in detail and I am not familiar with the Annoy library but I tried to replace \n`val_df = pd.read_csv('/kaggle/input/hotel-id-vector-extraction/validation.csv')`\nwith \n`val_df = pd.read_csv('/kaggle/input/hotel-id-vector-extraction/training.csv')`\nin your Hotel ID - Vector Validation notebook and it still get zero correct matches. I would expect to get them all correctly so there must be something wrong.\n\nI found one error in the Hotel ID - Vector Validation candidate finding part\nYou should replace this:\n```\nfor hotel, vs in enumerate(votes):\n\tif vs > max_votes:\n```\n\nwith \n```\nfor i, hotel in enumerate(votes.keys()):\n\tvs = votes[hotel]\n\tif vs > max_votes:\n```\nbecause in your case the enumerate returns index into the *hotel* variable and hotel_id (key from the map) into the *vs* variable. Then it get at least some correct for the training data and few for valid dataset. But still it's not great so it would be worth checking why it can't classify even the training data correctly.\n\nI can imagine few ways you could improve the solution:\n- You could train the model to classify the images first and then extract the features with the finetuned model. This should give the model a better chance to learn features important for the task.\n- The orientation of the images might be important in this case so you might want to identify the ones with incorrect orientation and rotate them accordingly. \n- And maybe there is a better way to compare the features. If i undestand it correctly you generate 320 features of size 8x8 for each image and then do knn search for each of the 8x8 points accross all features. Maybe it would be better to just unpack the 8x8 features into 64 values and do search on them or you could try to use something like PCA to reduce the dimension size.",
    "1752559": "Thank you, fixed that bug by changing it to votes.items(), which is what I meant to use. That improves it to get 10 of the validation set right.\n\nGood leads on training a classifier and fixing orientation, thank you.\n\nIf I'm understanding correctly, your suggested method here is what I'm doing. I'm extracting 64 (8x8) 320-dimensional feature vectors from each image, which are what are being searched.",
    "1752837": "Well my suggestion was more like doing 320 searches on 64 value feature vector so you would be looking for images with similar values in each feature or just do one knn on all feature values at once. But maybe it would lead to the same result not sure (it's just searching in different dimension I guess). But if you do 320 searches for each feature you could try to identify which feature was important (lead to most correct predictions) and visualize it for fun or give them weights in final voting.\n\nDid you try some other method than Annoy? When I tried to classify training data, use only the base_transform (no occlusions) and K=1 it still didn't get very good results. Even though for KNN I would expect the nearest neighbor would be the input image itself."
  },
  "source": "meta"
}