{
  "id": 163508,
  "title": "What are the constrains for generating embedding",
  "url": "/competitions/landmark-retrieval-2020/discussion/163508",
  "author_name": "",
  "post_date": "2020-07-02T10:29:43.172719600Z",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi Team,\n     i have a small doubt about Embedding generation. are there any constrains for the embedding?. like length(say 128), normalized to unit vector , etc.... pls clarify. </p>\n\n<p>thanks\nyuvaram</p>",
  "messages": [
    {
      "id": "912227",
      "postDate": "07/02/2020 10:29:43",
      "content": "<p>Hi Team,\n     i have a small doubt about Embedding generation. are there any constrains for the embedding?. like length(say 128), normalized to unit vector , etc.... pls clarify. </p>\n\n<p>thanks\nyuvaram</p>",
      "rawMarkdown": "Hi Team,\n     i have a small doubt about Embedding generation. are there any constrains for the embedding?. like length(say 128), normalized to unit vector , etc.... pls clarify. \n\nthanks\nyuvaram",
      "votes": null
    },
    {
      "id": "912922",
      "postDate": "07/02/2020 20:19:36",
      "content": "<p>The only constraint is that Euclidean distance is used for nearest neighbor computation. You can use any length, and it does not need to be normalized.</p>\n\n<p>Keep in mind, though, of another constraint, which is that the full feature extraction + nearest neighbor lookup should fit within the 9 hours budget (as mentioned in the <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/overview/code-requirements\">Code Requirements</a> page). So, an embedding that takes extremely long to compute, or a really large embedding that makes brute-force nearest-neighbor comparisons really slow, are not going to work.</p>",
      "rawMarkdown": "The only constraint is that Euclidean distance is used for nearest neighbor computation. You can use any length, and it does not need to be normalized.\n\nKeep in mind, though, of another constraint, which is that the full feature extraction + nearest neighbor lookup should fit within the 9 hours budget (as mentioned in the [Code Requirements](https://www.kaggle.com/c/landmark-retrieval-2020/overview/code-requirements) page). So, an embedding that takes extremely long to compute, or a really large embedding that makes brute-force nearest-neighbor comparisons really slow, are not going to work.",
      "votes": null
    },
    {
      "id": "913223",
      "postDate": "07/03/2020 04:33:04",
      "content": "<p>Thanks <a href=\"/andrefaraujo\">@andrefaraujo</a> </p>",
      "rawMarkdown": "Thanks @andrefaraujo",
      "votes": null
    },
    {
      "id": "913732",
      "postDate": "07/03/2020 11:30:15",
      "content": "<p>Do you do efficient NN via faiss-gpu, or is it just argmin on huge distance matrix?</p>",
      "rawMarkdown": "Do you do efficient NN via faiss-gpu, or is it just argmin on huge distance matrix?",
      "votes": null
    },
    {
      "id": "914079",
      "postDate": "07/03/2020 16:04:02",
      "content": "<p>We are using a brute-force nearest neighbor approach -- so more like the latter.</p>",
      "rawMarkdown": "We are using a brute-force nearest neighbor approach -- so more like the latter.",
      "votes": null
    },
    {
      "id": "915050",
      "postDate": "07/04/2020 12:55:53",
      "content": "<p>Well, faiss-gpu has brute force as well - that is what I meant. It is just that the implementation is very efficient - both for memory and computations.</p>",
      "rawMarkdown": "Well, faiss-gpu has brute force as well - that is what I meant. It is just that the implementation is very efficient - both for memory and computations.",
      "votes": null
    },
    {
      "id": "916497",
      "postDate": "07/05/2020 17:47:25",
      "content": "<p>Oh, I see. Anyway, we are not using faiss.</p>",
      "rawMarkdown": "Oh, I see. Anyway, we are not using faiss.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 912922,
      "author_name": "andrefaraujo",
      "author_url": "",
      "post_date": "07/02/2020 20:19:36",
      "content": "<p>The only constraint is that Euclidean distance is used for nearest neighbor computation. You can use any length, and it does not need to be normalized.</p>\n\n<p>Keep in mind, though, of another constraint, which is that the full feature extraction + nearest neighbor lookup should fit within the 9 hours budget (as mentioned in the <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/overview/code-requirements\">Code Requirements</a> page). So, an embedding that takes extremely long to compute, or a really large embedding that makes brute-force nearest-neighbor comparisons really slow, are not going to work.</p>",
      "votes": null,
      "replies": [
        {
          "id": 913223,
          "author_name": "yuvaramsingh",
          "author_url": "",
          "post_date": "07/03/2020 04:33:04",
          "content": "<p>Thanks <a href=\"/andrefaraujo\">@andrefaraujo</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 913732,
          "author_name": "oldufo",
          "author_url": "",
          "post_date": "07/03/2020 11:30:15",
          "content": "<p>Do you do efficient NN via faiss-gpu, or is it just argmin on huge distance matrix?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 914079,
          "author_name": "andrefaraujo",
          "author_url": "",
          "post_date": "07/03/2020 16:04:02",
          "content": "<p>We are using a brute-force nearest neighbor approach -- so more like the latter.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 915050,
          "author_name": "oldufo",
          "author_url": "",
          "post_date": "07/04/2020 12:55:53",
          "content": "<p>Well, faiss-gpu has brute force as well - that is what I meant. It is just that the implementation is very efficient - both for memory and computations.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 916497,
          "author_name": "andrefaraujo",
          "author_url": "",
          "post_date": "07/05/2020 17:47:25",
          "content": "<p>Oh, I see. Anyway, we are not using faiss.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "912227": "Hi Team,\n     i have a small doubt about Embedding generation. are there any constrains for the embedding?. like length(say 128), normalized to unit vector , etc.... pls clarify. \n\nthanks\nyuvaram",
    "912922": "The only constraint is that Euclidean distance is used for nearest neighbor computation. You can use any length, and it does not need to be normalized.\n\nKeep in mind, though, of another constraint, which is that the full feature extraction + nearest neighbor lookup should fit within the 9 hours budget (as mentioned in the [Code Requirements](https://www.kaggle.com/c/landmark-retrieval-2020/overview/code-requirements) page). So, an embedding that takes extremely long to compute, or a really large embedding that makes brute-force nearest-neighbor comparisons really slow, are not going to work.",
    "913223": "Thanks @andrefaraujo",
    "913732": "Do you do efficient NN via faiss-gpu, or is it just argmin on huge distance matrix?",
    "914079": "We are using a brute-force nearest neighbor approach -- so more like the latter.",
    "915050": "Well, faiss-gpu has brute force as well - that is what I meant. It is just that the implementation is very efficient - both for memory and computations.",
    "916497": "Oh, I see. Anyway, we are not using faiss."
  },
  "source": "meta"
}