{
  "id": 176697,
  "title": "Can we pre-compute embeddings? Are image ids the same in PB and PV sets?",
  "url": "/competitions/landmark-recognition-2020/discussion/176697",
  "author_name": "Chan Kha Vu",
  "post_date": "2020-08-23T03:10:27.080000",
  "votes": 6,
  "comment_count": 15,
  "views": 0,
  "content": "<p>From the bird-eyes view, the prediction pipeline looks as follows:</p>\n<ol>\n<li>First, compute global and local embedding vectors for both <strong>train</strong> and <strong>test</strong> sets.</li>\n<li>Using some weird <em>voodoo art of deep learning alchemy</em>, for each query in test set, find images from the train set that looks similar to the query image. Extract class id from there.</li>\n</ol>\n<p>Now, the question is:</p>\n<blockquote>\n  <p>Do we really have to squeeze step 1 into 12 hours limit? Or <strong>can we pre-compute the embeddings for train set</strong> and then apply them during prediction? To be able to do that, the <strong>image id</strong> should match in both <strong>public</strong> and <strong>private</strong> train sets.</p>\n</blockquote>\n<p>The competition rules only says:</p>\n<blockquote>\n  <p>To facilitate recognition-by-retrieval approaches, the private training set contains only a 100k subset of the total public training set. This 100k subset contains all of the training set images associated with the landmarks in the private test set.</p>\n</blockquote>\n<p>So, there's no guarantee that the image ids would match? Can we pre-compute image hashes than?</p>\n<hr>\n<p>Related questions on this topic (public/private consistency):</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/176319\" target=\"_blank\">How many unique landmarks are in the testing set?</a></li>\n<li><a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/175171\" target=\"_blank\">Are GLRec and GLRet datasets identical?</a></li>\n<li><a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/173111\" target=\"_blank\">What exactly is the 100k private training set?</a></li>\n<li><a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/174919\" target=\"_blank\">Can someone clarify the public and private data set for me?</a></li>\n</ul>",
  "messages": [
    {
      "id": 982067,
      "postDate": "2020-08-23T03:10:27.080Z",
      "content": "<p>From the bird-eyes view, the prediction pipeline looks as follows:</p>\n<ol>\n<li>First, compute global and local embedding vectors for both <strong>train</strong> and <strong>test</strong> sets.</li>\n<li>Using some weird <em>voodoo art of deep learning alchemy</em>, for each query in test set, find images from the train set that looks similar to the query image. Extract class id from there.</li>\n</ol>\n<p>Now, the question is:</p>\n<blockquote>\n  <p>Do we really have to squeeze step 1 into 12 hours limit? Or <strong>can we pre-compute the embeddings for train set</strong> and then apply them during prediction? To be able to do that, the <strong>image id</strong> should match in both <strong>public</strong> and <strong>private</strong> train sets.</p>\n</blockquote>\n<p>The competition rules only says:</p>\n<blockquote>\n  <p>To facilitate recognition-by-retrieval approaches, the private training set contains only a 100k subset of the total public training set. This 100k subset contains all of the training set images associated with the landmarks in the private test set.</p>\n</blockquote>\n<p>So, there's no guarantee that the image ids would match? Can we pre-compute image hashes than?</p>\n<hr>\n<p>Related questions on this topic (public/private consistency):</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/176319\" target=\"_blank\">How many unique landmarks are in the testing set?</a></li>\n<li><a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/175171\" target=\"_blank\">Are GLRec and GLRet datasets identical?</a></li>\n<li><a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/173111\" target=\"_blank\">What exactly is the 100k private training set?</a></li>\n<li><a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/174919\" target=\"_blank\">Can someone clarify the public and private data set for me?</a></li>\n</ul>",
      "rawMarkdown": "From the bird-eyes view, the prediction pipeline looks as follows:\n\n  1. First, compute global and local embedding vectors for both **train** and **test** sets.\n  2. Using some weird *voodoo art of deep learning alchemy*, for each query in test set, find images from the train set that looks similar to the query image. Extract class id from there.\n\nNow, the question is:\n\n> Do we really have to squeeze step 1 into 12 hours limit? Or **can we pre-compute the embeddings for train set** and then apply them during prediction? To be able to do that, the **image id** should match in both **public** and **private** train sets.\n\nThe competition rules only says:\n> To facilitate recognition-by-retrieval approaches, the private training set contains only a 100k subset of the total public training set. This 100k subset contains all of the training set images associated with the landmarks in the private test set.\n\nSo, there's no guarantee that the image ids would match? Can we pre-compute image hashes than?\n\n-----------------------------------------------------------------\n\nRelated questions on this topic (public/private consistency):\n\n* [How many unique landmarks are in the testing set?](https://www.kaggle.com/c/landmark-recognition-2020/discussion/176319)\n* [Are GLRec and GLRet datasets identical?](https://www.kaggle.com/c/landmark-recognition-2020/discussion/175171)\n* [What exactly is the 100k private training set?](https://www.kaggle.com/c/landmark-recognition-2020/discussion/173111)\n* [Can someone clarify the public and private data set for me?](https://www.kaggle.com/c/landmark-recognition-2020/discussion/174919)",
      "votes": 6
    },
    {
      "id": 993318,
      "postDate": "2020-08-31T20:43:43.860Z",
      "content": "<p>It would be nice if an organizer could respond to this thread, this is quite important. <br>\nWe know we can upload the whole training set as external data, but can we upload pre-computed embeddings? Be it for each training sample, class centers or anything else related to the train data? <a href=\"https://www.kaggle.com/andrefaraujo\" target=\"_blank\">@andrefaraujo</a> <a href=\"https://www.kaggle.com/camaskew\" target=\"_blank\">@camaskew</a> <a href=\"https://www.kaggle.com/tobwey\" target=\"_blank\">@tobwey</a> <br>\nThanks!</p>",
      "rawMarkdown": "It would be nice if an organizer could respond to this thread, this is quite important. \nWe know we can upload the whole training set as external data, but can we upload pre-computed embeddings? Be it for each training sample, class centers or anything else related to the train data? @andrefaraujo @camaskew @tobwey \nThanks!",
      "votes": 1,
      "replies": [
        {
          "id": 993334,
          "postDate": "2020-08-31T21:03:58.703Z",
          "content": "<p>any easier way to use entire training set apart from splitting it into 5 part, 20GB kaggle datasets?</p>",
          "rawMarkdown": "any easier way to use entire training set apart from splitting it into 5 part, 20GB kaggle datasets?"
        },
        {
          "id": 993366,
          "postDate": "2020-08-31T21:42:14.690Z",
          "content": "<p>No ideia haha</p>",
          "rawMarkdown": "No ideia haha"
        },
        {
          "id": 993747,
          "postDate": "2020-09-01T05:59:23.190Z",
          "content": "<p>Yes, you can upload pre-computed embeddings. Similar to what <a href=\"https://www.kaggle.com/skrish13\" target=\"_blank\">@skrish13</a> mentioned below, the pre-computed embeddings could be seen as the parameters of a kNN classifier, so that's totally fine. Similar answer for class centers / other things related to the train data -- they could be seen as part of your recognition model.</p>",
          "rawMarkdown": "Yes, you can upload pre-computed embeddings. Similar to what @skrish13 mentioned below, the pre-computed embeddings could be seen as the parameters of a kNN classifier, so that's totally fine. Similar answer for class centers / other things related to the train data -- they could be seen as part of your recognition model.",
          "votes": 5
        },
        {
          "id": 994086,
          "postDate": "2020-09-01T11:04:31.200Z",
          "content": "<p>Thanks a lot!</p>",
          "rawMarkdown": "Thanks a lot!"
        },
        {
          "id": 994087,
          "postDate": "2020-09-01T11:04:31.237Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 994107,
          "postDate": "2020-09-01T11:19:40.680Z",
          "content": "<p><a href=\"https://www.kaggle.com/andrefaraujo\" target=\"_blank\">@andrefaraujo</a> the images indexed the same in public and private training sets? For example, will <code>0018aa4b92532b77.jpeg</code> in train set be the same image (down to bits) as <code>0018aa4b92532b77.jpeg</code> in private set?</p>",
          "rawMarkdown": "@andrefaraujo the images indexed the same in public and private training sets? For example, will `0018aa4b92532b77.jpeg` in train set be the same image (down to bits) as `0018aa4b92532b77.jpeg` in private set?"
        },
        {
          "id": 997097,
          "postDate": "2020-09-03T18:33:28.030Z",
          "content": "<p>Yes, the images should be identical.</p>",
          "rawMarkdown": "Yes, the images should be identical.",
          "votes": 2
        }
      ]
    },
    {
      "id": 982169,
      "postDate": "2020-08-23T05:46:06.830Z",
      "content": "<p>I've been wondering exactly this and my experiments suggest they do not match, but I'm unsure and would appreciate clarification from the organizers.</p>",
      "rawMarkdown": "I've been wondering exactly this and my experiments suggest they do not match, but I'm unsure and would appreciate clarification from the organizers.",
      "votes": 2,
      "replies": [
        {
          "id": 984496,
          "postDate": "2020-08-25T05:48:48.093Z",
          "content": "<p>Can you please describe your experiment? Thanks a lot for the answer btw!</p>",
          "rawMarkdown": "Can you please describe your experiment? Thanks a lot for the answer btw!"
        },
        {
          "id": 985505,
          "postDate": "2020-08-25T19:01:35.213Z",
          "content": "<p>I tried to calculate an md5sum over a given id locally and in the notebook. I throw an exception and crash the notebook during rerun if it doesn't match the string I entered. If it matches, I submit the sample_submission.csv. This way you can tell they have changed something because the submission will fail.</p>",
          "rawMarkdown": "I tried to calculate an md5sum over a given id locally and in the notebook. I throw an exception and crash the notebook during rerun if it doesn't match the string I entered. If it matches, I submit the sample_submission.csv. This way you can tell they have changed something because the submission will fail.",
          "votes": 1
        }
      ]
    },
    {
      "id": 983671,
      "postDate": "2020-08-24T13:49:33.520Z",
      "content": "<p>Code Requirements says that</p>\n<ul>\n<li>Freely &amp; publicly available external data is allowed, including pre-trained models</li>\n</ul>\n<p>So, I'm not sure whether you can use your own pre-computed embeddings if it's not publicly available. I'd like some more clarification on that too.</p>",
      "rawMarkdown": "Code Requirements says that\n\n- Freely & publicly available external data is allowed, including pre-trained models\n\nSo, I'm not sure whether you can use your own pre-computed embeddings if it's not publicly available. I'd like some more clarification on that too.",
      "replies": [
        {
          "id": 989202,
          "postDate": "2020-08-28T16:03:33.260Z",
          "content": "<p>So if I train a model and save it in one script (s1)  and I add it as data onto another script file  (s2).. is it considered as external data? I am submitting s2 for scoring. </p>",
          "rawMarkdown": "So if I train a model and save it in one script (s1)  and I add it as data onto another script file  (s2).. is it considered as external data? I am submitting s2 for scoring. \n"
        }
      ]
    },
    {
      "id": 991081,
      "postDate": "2020-08-30T06:09:32.610Z",
      "content": "<p>If you can use a trained classifier then I'd assume you can use pre-computed embeddings for sure.</p>",
      "rawMarkdown": "If you can use a trained classifier then I'd assume you can use pre-computed embeddings for sure."
    },
    {
      "id": 982971,
      "postDate": "2020-08-23T22:33:54.563Z",
      "content": "<p>If i undestand correctly you can use another training set, so you could calculate train embeddings previously and attach them as an external dataset.</p>\n<p>\"You may still attach the full training set as an external data set if you wish.\"</p>",
      "rawMarkdown": "If i undestand correctly you can use another training set, so you could calculate train embeddings previously and attach them as an external dataset.\n\n\"You may still attach the full training set as an external data set if you wish.\""
    }
  ],
  "comments": [
    {
      "id": 993318,
      "author_name": "Eduardo Rocha de Andrade",
      "author_url": "",
      "post_date": "2020-08-31T20:43:43.860000",
      "content": "<p>It would be nice if an organizer could respond to this thread, this is quite important. <br>\nWe know we can upload the whole training set as external data, but can we upload pre-computed embeddings? Be it for each training sample, class centers or anything else related to the train data? <a href=\"https://www.kaggle.com/andrefaraujo\" target=\"_blank\">@andrefaraujo</a> <a href=\"https://www.kaggle.com/camaskew\" target=\"_blank\">@camaskew</a> <a href=\"https://www.kaggle.com/tobwey\" target=\"_blank\">@tobwey</a> <br>\nThanks!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 993334,
          "author_name": "hirviö",
          "author_url": "",
          "post_date": "2020-08-31T21:03:58.703000",
          "content": "<p>any easier way to use entire training set apart from splitting it into 5 part, 20GB kaggle datasets?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 993366,
          "author_name": "Eduardo Rocha de Andrade",
          "author_url": "",
          "post_date": "2020-08-31T21:42:14.690000",
          "content": "<p>No ideia haha</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 993747,
          "author_name": "Andre Araujo",
          "author_url": "",
          "post_date": "2020-09-01T05:59:23.190000",
          "content": "<p>Yes, you can upload pre-computed embeddings. Similar to what <a href=\"https://www.kaggle.com/skrish13\" target=\"_blank\">@skrish13</a> mentioned below, the pre-computed embeddings could be seen as the parameters of a kNN classifier, so that's totally fine. Similar answer for class centers / other things related to the train data -- they could be seen as part of your recognition model.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 994086,
          "author_name": "Eduardo Rocha de Andrade",
          "author_url": "",
          "post_date": "2020-09-01T11:04:31.200000",
          "content": "<p>Thanks a lot!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 994087,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-09-01T11:04:31.237000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 994107,
          "author_name": "Chan Kha Vu",
          "author_url": "",
          "post_date": "2020-09-01T11:19:40.680000",
          "content": "<p><a href=\"https://www.kaggle.com/andrefaraujo\" target=\"_blank\">@andrefaraujo</a> the images indexed the same in public and private training sets? For example, will <code>0018aa4b92532b77.jpeg</code> in train set be the same image (down to bits) as <code>0018aa4b92532b77.jpeg</code> in private set?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 997097,
          "author_name": "Andre Araujo",
          "author_url": "",
          "post_date": "2020-09-03T18:33:28.030000",
          "content": "<p>Yes, the images should be identical.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 982169,
      "author_name": "Usmann Khan",
      "author_url": "",
      "post_date": "2020-08-23T05:46:06.830000",
      "content": "<p>I've been wondering exactly this and my experiments suggest they do not match, but I'm unsure and would appreciate clarification from the organizers.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 984496,
          "author_name": "Chan Kha Vu",
          "author_url": "",
          "post_date": "2020-08-25T05:48:48.093000",
          "content": "<p>Can you please describe your experiment? Thanks a lot for the answer btw!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 985505,
          "author_name": "Usmann Khan",
          "author_url": "",
          "post_date": "2020-08-25T19:01:35.213000",
          "content": "<p>I tried to calculate an md5sum over a given id locally and in the notebook. I throw an exception and crash the notebook during rerun if it doesn't match the string I entered. If it matches, I submit the sample_submission.csv. This way you can tell they have changed something because the submission will fail.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 983671,
      "author_name": "Yusuf Büyükdağ",
      "author_url": "",
      "post_date": "2020-08-24T13:49:33.520000",
      "content": "<p>Code Requirements says that</p>\n<ul>\n<li>Freely &amp; publicly available external data is allowed, including pre-trained models</li>\n</ul>\n<p>So, I'm not sure whether you can use your own pre-computed embeddings if it's not publicly available. I'd like some more clarification on that too.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 989202,
          "author_name": "_CA℟L_",
          "author_url": "",
          "post_date": "2020-08-28T16:03:33.260000",
          "content": "<p>So if I train a model and save it in one script (s1)  and I add it as data onto another script file  (s2).. is it considered as external data? I am submitting s2 for scoring. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 991081,
      "author_name": "hirviö",
      "author_url": "",
      "post_date": "2020-08-30T06:09:32.610000",
      "content": "<p>If you can use a trained classifier then I'd assume you can use pre-computed embeddings for sure.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 982971,
      "author_name": "Martin Kovacevic Buvinic",
      "author_url": "",
      "post_date": "2020-08-23T22:33:54.563000",
      "content": "<p>If i undestand correctly you can use another training set, so you could calculate train embeddings previously and attach them as an external dataset.</p>\n<p>\"You may still attach the full training set as an external data set if you wish.\"</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "982067": "From the bird-eyes view, the prediction pipeline looks as follows:\n\n  1. First, compute global and local embedding vectors for both **train** and **test** sets.\n  2. Using some weird *voodoo art of deep learning alchemy*, for each query in test set, find images from the train set that looks similar to the query image. Extract class id from there.\n\nNow, the question is:\n\n> Do we really have to squeeze step 1 into 12 hours limit? Or **can we pre-compute the embeddings for train set** and then apply them during prediction? To be able to do that, the **image id** should match in both **public** and **private** train sets.\n\nThe competition rules only says:\n> To facilitate recognition-by-retrieval approaches, the private training set contains only a 100k subset of the total public training set. This 100k subset contains all of the training set images associated with the landmarks in the private test set.\n\nSo, there's no guarantee that the image ids would match? Can we pre-compute image hashes than?\n\n-----------------------------------------------------------------\n\nRelated questions on this topic (public/private consistency):\n\n* [How many unique landmarks are in the testing set?](https://www.kaggle.com/c/landmark-recognition-2020/discussion/176319)\n* [Are GLRec and GLRet datasets identical?](https://www.kaggle.com/c/landmark-recognition-2020/discussion/175171)\n* [What exactly is the 100k private training set?](https://www.kaggle.com/c/landmark-recognition-2020/discussion/173111)\n* [Can someone clarify the public and private data set for me?](https://www.kaggle.com/c/landmark-recognition-2020/discussion/174919)",
    "993318": "It would be nice if an organizer could respond to this thread, this is quite important. \nWe know we can upload the whole training set as external data, but can we upload pre-computed embeddings? Be it for each training sample, class centers or anything else related to the train data? @andrefaraujo @camaskew @tobwey \nThanks!",
    "982169": "I've been wondering exactly this and my experiments suggest they do not match, but I'm unsure and would appreciate clarification from the organizers.",
    "983671": "Code Requirements says that\n\n- Freely & publicly available external data is allowed, including pre-trained models\n\nSo, I'm not sure whether you can use your own pre-computed embeddings if it's not publicly available. I'd like some more clarification on that too.",
    "991081": "If you can use a trained classifier then I'd assume you can use pre-computed embeddings for sure.",
    "982971": "If i undestand correctly you can use another training set, so you could calculate train embeddings previously and attach them as an external dataset.\n\n\"You may still attach the full training set as an external data set if you wish.\""
  }
}