{
  "id": 317609,
  "title": "Concatenation issue during the ensembling.",
  "url": "/competitions/happy-whale-and-dolphin/discussion/317609",
  "author_name": "York G",
  "post_date": "2022-04-08T02:47:12.031000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I am using the concatenation strategy on my models ensemble. But after doing that, LB score dropped which was out of my expectation. <br>\nI think there should be some mistakes in my processing, but can't figure it out. </p>\n<p>The below steps are what I did, is  there anything I missed? <br>\nThank you for the advice in advance. </p>\n<p>P.S. All the model trained on the same fold, so the shapes of embedding in each model are same.  </p>\n<ol>\n<li>Getting the embedding from model evaluation. </li>\n</ol>\n<pre><code>train_image_names, train_embeddings, train_targets = get_embeddings(model, train_dl, encoder, stage=\"train\")\n\nval_image_names, val_embeddings, val_targets = get_embeddings(model, val_dl, encoder, stage=\"val\")\n\ntest_image_names, test_embeddings, test_targets = get_embeddings(model, test_dl, encoder, stage=\"test\")\n</code></pre>\n<p>for example, The shape of val_embedding result is (8315, 768).</p>\n<ol>\n<li><p>Doing the normalization on the embedding results:<br>\n<code>tensor_train_embedding = F.normalize(torch.tensor(train_embeddings), dim=1)\ntensor_val_embedding = F.normalize(torch.tensor(val_embeddings), dim=1)\ntensor_test_embedding = F.normalize(torch.tensor(test_embeddings), dim=1)</code></p></li>\n<li><p>Concatenation <br>\n`<br>\ntt_train_embedding = np.concatenate([tt_train_embedding, tensor_train_embedding.detach().numpy()], axis=1)</p></li>\n</ol>\n<p>tt_val_embedding = np.concatenate([tt_val_embedding, tensor_val_embedding.detach().numpy()], axis=1)</p>\n<p>tt_test_embedding = np.concatenate([tt_test_embedding, tensor_test_embedding.detach().numpy()], axis=1)<br>\n`</p>\n<ol>\n<li>Doing the index and distance calculation:<br>\n` train_embeddings = tt_train_embedding<br>\nval_embeddings = tt_val_embedding<br>\ntest_embeddings = tt_test_embedding<br>\nmodel.embedding_size *= len(pretrained_list)</li>\n</ol>\n<p>D, I = wpi.create_and_search_index(model.embedding_size, train_embeddings, val_embeddings, k)<br>\n… …<br>\n`<br>\nThe len(pretrained_list) is how many models used in emsemble.</p>",
  "messages": [
    {
      "id": 1748808,
      "postDate": "2022-04-08T02:47:12.033Z",
      "content": "<p>I am using the concatenation strategy on my models ensemble. But after doing that, LB score dropped which was out of my expectation. <br>\nI think there should be some mistakes in my processing, but can't figure it out. </p>\n<p>The below steps are what I did, is  there anything I missed? <br>\nThank you for the advice in advance. </p>\n<p>P.S. All the model trained on the same fold, so the shapes of embedding in each model are same.  </p>\n<ol>\n<li>Getting the embedding from model evaluation. </li>\n</ol>\n<pre><code>train_image_names, train_embeddings, train_targets = get_embeddings(model, train_dl, encoder, stage=\"train\")\n\nval_image_names, val_embeddings, val_targets = get_embeddings(model, val_dl, encoder, stage=\"val\")\n\ntest_image_names, test_embeddings, test_targets = get_embeddings(model, test_dl, encoder, stage=\"test\")\n</code></pre>\n<p>for example, The shape of val_embedding result is (8315, 768).</p>\n<ol>\n<li><p>Doing the normalization on the embedding results:<br>\n<code>tensor_train_embedding = F.normalize(torch.tensor(train_embeddings), dim=1)\ntensor_val_embedding = F.normalize(torch.tensor(val_embeddings), dim=1)\ntensor_test_embedding = F.normalize(torch.tensor(test_embeddings), dim=1)</code></p></li>\n<li><p>Concatenation <br>\n`<br>\ntt_train_embedding = np.concatenate([tt_train_embedding, tensor_train_embedding.detach().numpy()], axis=1)</p></li>\n</ol>\n<p>tt_val_embedding = np.concatenate([tt_val_embedding, tensor_val_embedding.detach().numpy()], axis=1)</p>\n<p>tt_test_embedding = np.concatenate([tt_test_embedding, tensor_test_embedding.detach().numpy()], axis=1)<br>\n`</p>\n<ol>\n<li>Doing the index and distance calculation:<br>\n` train_embeddings = tt_train_embedding<br>\nval_embeddings = tt_val_embedding<br>\ntest_embeddings = tt_test_embedding<br>\nmodel.embedding_size *= len(pretrained_list)</li>\n</ol>\n<p>D, I = wpi.create_and_search_index(model.embedding_size, train_embeddings, val_embeddings, k)<br>\n… …<br>\n`<br>\nThe len(pretrained_list) is how many models used in emsemble.</p>",
      "rawMarkdown": "I am using the concatenation strategy on my models ensemble. But after doing that, LB score dropped which was out of my expectation. \nI think there should be some mistakes in my processing, but can't figure it out. \n\nThe below steps are what I did, is  there anything I missed? \nThank you for the advice in advance. \n\nP.S. All the model trained on the same fold, so the shapes of embedding in each model are same.  \n\n1. Getting the embedding from model evaluation. \n```\ntrain_image_names, train_embeddings, train_targets = get_embeddings(model, train_dl, encoder, stage=\"train\")\n\nval_image_names, val_embeddings, val_targets = get_embeddings(model, val_dl, encoder, stage=\"val\")\n\ntest_image_names, test_embeddings, test_targets = get_embeddings(model, test_dl, encoder, stage=\"test\")\n```\nfor example, The shape of val_embedding result is (8315, 768).\n\n2. Doing the normalization on the embedding results:\n` tensor_train_embedding = F.normalize(torch.tensor(train_embeddings), dim=1)\ntensor_val_embedding = F.normalize(torch.tensor(val_embeddings), dim=1)\ntensor_test_embedding = F.normalize(torch.tensor(test_embeddings), dim=1)`\n\n3. Concatenation \n`\ntt_train_embedding = np.concatenate([tt_train_embedding, tensor_train_embedding.detach().numpy()], axis=1)\n\ntt_val_embedding = np.concatenate([tt_val_embedding, tensor_val_embedding.detach().numpy()], axis=1)\n\ntt_test_embedding = np.concatenate([tt_test_embedding, tensor_test_embedding.detach().numpy()], axis=1)\n`\n\n4. Doing the index and distance calculation:\n` train_embeddings = tt_train_embedding\n    val_embeddings = tt_val_embedding\n    test_embeddings = tt_test_embedding\n    model.embedding_size *= len(pretrained_list)\n\nD, I = wpi.create_and_search_index(model.embedding_size, train_embeddings, val_embeddings, k)\n... ...\n`\nThe len(pretrained_list) is how many models used in emsemble.\n ",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1748808": "I am using the concatenation strategy on my models ensemble. But after doing that, LB score dropped which was out of my expectation. \nI think there should be some mistakes in my processing, but can't figure it out. \n\nThe below steps are what I did, is  there anything I missed? \nThank you for the advice in advance. \n\nP.S. All the model trained on the same fold, so the shapes of embedding in each model are same.  \n\n1. Getting the embedding from model evaluation. \n```\ntrain_image_names, train_embeddings, train_targets = get_embeddings(model, train_dl, encoder, stage=\"train\")\n\nval_image_names, val_embeddings, val_targets = get_embeddings(model, val_dl, encoder, stage=\"val\")\n\ntest_image_names, test_embeddings, test_targets = get_embeddings(model, test_dl, encoder, stage=\"test\")\n```\nfor example, The shape of val_embedding result is (8315, 768).\n\n2. Doing the normalization on the embedding results:\n` tensor_train_embedding = F.normalize(torch.tensor(train_embeddings), dim=1)\ntensor_val_embedding = F.normalize(torch.tensor(val_embeddings), dim=1)\ntensor_test_embedding = F.normalize(torch.tensor(test_embeddings), dim=1)`\n\n3. Concatenation \n`\ntt_train_embedding = np.concatenate([tt_train_embedding, tensor_train_embedding.detach().numpy()], axis=1)\n\ntt_val_embedding = np.concatenate([tt_val_embedding, tensor_val_embedding.detach().numpy()], axis=1)\n\ntt_test_embedding = np.concatenate([tt_test_embedding, tensor_test_embedding.detach().numpy()], axis=1)\n`\n\n4. Doing the index and distance calculation:\n` train_embeddings = tt_train_embedding\n    val_embeddings = tt_val_embedding\n    test_embeddings = tt_test_embedding\n    model.embedding_size *= len(pretrained_list)\n\nD, I = wpi.create_and_search_index(model.embedding_size, train_embeddings, val_embeddings, k)\n... ...\n`\nThe len(pretrained_list) is how many models used in emsemble.\n "
  }
}