{
  "id": 277601,
  "title": "Matching multilingual BERT embeddings with image embeddings",
  "url": "/competitions/wikipedia-image-caption/discussion/277601",
  "author_name": "",
  "post_date": "2021-10-10T09:09:11.097871Z",
  "votes": 10,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi everyone!</p>\n<p>I wanted to share a strategy that I would have followed if I currently had more time and the necessary compute resources. I'd be happy to hear your thoughts about the strategy (maybe it's even the completely standard approach?) and maybe someone with more time and more compute resources wants to try it and can report his results.</p>\n<p>So here is my idea:<br>\n<strong>Step 1 (Word embeddings):</strong><br>\nUse the bert-base-multilingual-uncased model to obtain embeddings for the captions (e.g. \"Pariser Kanonen\" is tokenized into ['[CLS]', 'pariser', 'kanon', '##en', '[SEP]'] and the BERT model then yields for each of these 5 tokens an embedding, e.g. of dimension 768).</p>\n<p>For more details how this works, I recommend the following blog post:<br>\n<a href=\"url\" target=\"_blank\">https://mccormickml.com/2019/05/14/BERT-word-embeddings-tutorial/</a><br>\nFurther, you may use the code provided in the end of this post.</p>\n<p><strong>Step 2 (Create Database):</strong><br>\nCreate a database like the following:</p>\n<table>\n<thead>\n<tr>\n<th>Input 1</th>\n<th>Input 2</th>\n<th>Target</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Img_emb_PariserKanon</td>\n<td>BERT_emb_[CLS]</td>\n<td>1</td>\n</tr>\n<tr>\n<td>Img_emb_PariserKanon</td>\n<td>BERT_emb_pariser</td>\n<td>1</td>\n</tr>\n<tr>\n<td>Img_emb_PariserKanon</td>\n<td>BERT_emb_kanon</td>\n<td>1</td>\n</tr>\n<tr>\n<td>Img_emb_PariserKanon</td>\n<td>BERT_emb_##en</td>\n<td>1</td>\n</tr>\n<tr>\n<td>Img_emb_PariserKanon</td>\n<td>BERT_emb_[SEP]</td>\n<td>1</td>\n</tr>\n<tr>\n<td>Img_emb_PariserKanon</td>\n<td>BERT_emb_[word from some other caption]_1</td>\n<td>0</td>\n</tr>\n<tr>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n</tr>\n<tr>\n<td>Img_emb_PariserKanon</td>\n<td>BERT_emb_[word from some other caption]_kk</td>\n<td>0</td>\n</tr>\n</tbody>\n</table>\n<p>Where kk is some parameter, e.g. 20, for \"negative sampling\" (compare to word2Vec ideas, see e.g. this great blog post <a href=\"url\" target=\"_blank\">https://jalammar.github.io/illustrated-word2vec/</a> ). And the embeddings of all other images and captions are inserted analogously.</p>\n<p>Alternatively (to save space and compute ressources) one could try to only use the BERT_emb_[CLS] as these should contain information about the entire caption.</p>\n<p><strong>Step 3 (Train model):</strong><br>\nTrain a MLP on this database.</p>\n<p>Alternatively (to save compute ressources) one could build a model f that maps the BERT embeddings to R^2048, and then one simply computes the sigmoid of the scalar product (Img_emb,f(BERT_emb)).  </p>\n<p><strong>Step 4 (Inference):</strong><br>\nWhen we want to find the caption for some test image, we take the embedding of the image and match it to each test caption (e.g. compute <br>\n1/N * (MLP(Img_emb,BERT_emb_word1) + … + MLP(Img_emb,BERT_emb_wordN)) ) and predict the captions that match best.</p>\n<p>What do you think of this strategy?</p>\n<p>Best<br>\nTobias</p>\n<hr>\n<p>Code to get BERT embeddings:</p>\n<pre><code>import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport datatable as dt\nimport torch \nfrom transformers import BertTokenizer, BertModel\n\ndf = dt.fread('../input/wikipedia-image-caption/train-00000-of-00005.tsv')\n\ntokenizer = BertTokenizer.from_pretrained('bert-base-multilingual-uncased')\nmodel = BertModel.from_pretrained(\"bert-base-multilingual-uncased\", output_hidden_states = True)\nmodel.eval()\n\nntest = 10\ncapts = [capt.replace(' [SEP]','').replace('\\n',' ') for capt in df[:ntest,'caption_title_and_reference_description'].to_list()[0]]\nwith torch.no_grad():\n    for capt in capts:\n        print(capt)\n        tokenized = tokenizer.tokenize(\"[CLS] \" + capt + \" [SEP]\")\n        print(tokenized)\n        indexed = tokenizer.convert_tokens_to_ids(tokenized)\n        print(indexed)\n        segments = [1] * len(tokenized) #as a single sentence is fed to Bert, just put 1 for each token\n\n        tokens = torch.tensor([indexed])\n        segments = torch.tensor([segments])\n\n        outputs = model(tokens, segments)\n        hidden_states = outputs[2]\n        # tuple to tensor, get rid of batch dim, permute to #tokens x 13 hidden states x 768 model dim\n        token_embeddings = torch.stack(hidden_states, dim=0).squeeze(dim=1).permute(1,0,2)\n        # concatenate last 4 hiden states\n        token_embeddings = torch.flatten(token_embeddings[:,-4:],start_dim=1)\n        print(token_embeddings.size())\n</code></pre>",
  "messages": [
    {
      "id": "1540223",
      "postDate": "10/10/2021 09:09:11",
      "content": "<p>Hi everyone!</p>\n<p>I wanted to share a strategy that I would have followed if I currently had more time and the necessary compute resources. I'd be happy to hear your thoughts about the strategy (maybe it's even the completely standard approach?) and maybe someone with more time and more compute resources wants to try it and can report his results.</p>\n<p>So here is my idea:<br>\n<strong>Step 1 (Word embeddings):</strong><br>\nUse the bert-base-multilingual-uncased model to obtain embeddings for the captions (e.g. \"Pariser Kanonen\" is tokenized into ['[CLS]', 'pariser', 'kanon', '##en', '[SEP]'] and the BERT model then yields for each of these 5 tokens an embedding, e.g. of dimension 768).</p>\n<p>For more details how this works, I recommend the following blog post:<br>\n<a href=\"url\" target=\"_blank\">https://mccormickml.com/2019/05/14/BERT-word-embeddings-tutorial/</a><br>\nFurther, you may use the code provided in the end of this post.</p>\n<p><strong>Step 2 (Create Database):</strong><br>\nCreate a database like the following:</p>\n<table>\n<thead>\n<tr>\n<th>Input 1</th>\n<th>Input 2</th>\n<th>Target</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Img_emb_PariserKanon</td>\n<td>BERT_emb_[CLS]</td>\n<td>1</td>\n</tr>\n<tr>\n<td>Img_emb_PariserKanon</td>\n<td>BERT_emb_pariser</td>\n<td>1</td>\n</tr>\n<tr>\n<td>Img_emb_PariserKanon</td>\n<td>BERT_emb_kanon</td>\n<td>1</td>\n</tr>\n<tr>\n<td>Img_emb_PariserKanon</td>\n<td>BERT_emb_##en</td>\n<td>1</td>\n</tr>\n<tr>\n<td>Img_emb_PariserKanon</td>\n<td>BERT_emb_[SEP]</td>\n<td>1</td>\n</tr>\n<tr>\n<td>Img_emb_PariserKanon</td>\n<td>BERT_emb_[word from some other caption]_1</td>\n<td>0</td>\n</tr>\n<tr>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n</tr>\n<tr>\n<td>Img_emb_PariserKanon</td>\n<td>BERT_emb_[word from some other caption]_kk</td>\n<td>0</td>\n</tr>\n</tbody>\n</table>\n<p>Where kk is some parameter, e.g. 20, for \"negative sampling\" (compare to word2Vec ideas, see e.g. this great blog post <a href=\"url\" target=\"_blank\">https://jalammar.github.io/illustrated-word2vec/</a> ). And the embeddings of all other images and captions are inserted analogously.</p>\n<p>Alternatively (to save space and compute ressources) one could try to only use the BERT_emb_[CLS] as these should contain information about the entire caption.</p>\n<p><strong>Step 3 (Train model):</strong><br>\nTrain a MLP on this database.</p>\n<p>Alternatively (to save compute ressources) one could build a model f that maps the BERT embeddings to R^2048, and then one simply computes the sigmoid of the scalar product (Img_emb,f(BERT_emb)).  </p>\n<p><strong>Step 4 (Inference):</strong><br>\nWhen we want to find the caption for some test image, we take the embedding of the image and match it to each test caption (e.g. compute <br>\n1/N * (MLP(Img_emb,BERT_emb_word1) + … + MLP(Img_emb,BERT_emb_wordN)) ) and predict the captions that match best.</p>\n<p>What do you think of this strategy?</p>\n<p>Best<br>\nTobias</p>\n<hr>\n<p>Code to get BERT embeddings:</p>\n<pre><code>import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport datatable as dt\nimport torch \nfrom transformers import BertTokenizer, BertModel\n\ndf = dt.fread('../input/wikipedia-image-caption/train-00000-of-00005.tsv')\n\ntokenizer = BertTokenizer.from_pretrained('bert-base-multilingual-uncased')\nmodel = BertModel.from_pretrained(\"bert-base-multilingual-uncased\", output_hidden_states = True)\nmodel.eval()\n\nntest = 10\ncapts = [capt.replace(' [SEP]','').replace('\\n',' ') for capt in df[:ntest,'caption_title_and_reference_description'].to_list()[0]]\nwith torch.no_grad():\n    for capt in capts:\n        print(capt)\n        tokenized = tokenizer.tokenize(\"[CLS] \" + capt + \" [SEP]\")\n        print(tokenized)\n        indexed = tokenizer.convert_tokens_to_ids(tokenized)\n        print(indexed)\n        segments = [1] * len(tokenized) #as a single sentence is fed to Bert, just put 1 for each token\n\n        tokens = torch.tensor([indexed])\n        segments = torch.tensor([segments])\n\n        outputs = model(tokens, segments)\n        hidden_states = outputs[2]\n        # tuple to tensor, get rid of batch dim, permute to #tokens x 13 hidden states x 768 model dim\n        token_embeddings = torch.stack(hidden_states, dim=0).squeeze(dim=1).permute(1,0,2)\n        # concatenate last 4 hiden states\n        token_embeddings = torch.flatten(token_embeddings[:,-4:],start_dim=1)\n        print(token_embeddings.size())\n</code></pre>",
      "rawMarkdown": "Hi everyone!\n\nI wanted to share a strategy that I would have followed if I currently had more time and the necessary compute resources. I'd be happy to hear your thoughts about the strategy (maybe it's even the completely standard approach?) and maybe someone with more time and more compute resources wants to try it and can report his results.\n\nSo here is my idea:\n**Step 1 (Word embeddings):**\nUse the bert-base-multilingual-uncased model to obtain embeddings for the captions (e.g. \"Pariser Kanonen\" is tokenized into ['[CLS]', 'pariser', 'kanon', '##en', '[SEP]'] and the BERT model then yields for each of these 5 tokens an embedding, e.g. of dimension 768).\n\nFor more details how this works, I recommend the following blog post:\n[https://mccormickml.com/2019/05/14/BERT-word-embeddings-tutorial/](url)\nFurther, you may use the code provided in the end of this post.\n\n**Step 2 (Create Database):**\nCreate a database like the following:\n| Input 1 | Input 2 | Target |\n| --- | --- | --- |\n| Img_emb_PariserKanon  | BERT_emb_[CLS] | 1 |\n| Img_emb_PariserKanon  | BERT_emb_pariser | 1 |\n| Img_emb_PariserKanon  | BERT_emb_kanon | 1 |\n| Img_emb_PariserKanon  | BERT_emb_##en | 1 |\n| Img_emb_PariserKanon  | BERT_emb_[SEP] | 1 |\n| Img_emb_PariserKanon  | BERT_emb_[word from some other caption]_1 | 0 |\n| ... | ... | ... |\n| Img_emb_PariserKanon  | BERT_emb_[word from some other caption]_kk | 0 |\n\nWhere kk is some parameter, e.g. 20, for \"negative sampling\" (compare to word2Vec ideas, see e.g. this great blog post [https://jalammar.github.io/illustrated-word2vec/](url) ). And the embeddings of all other images and captions are inserted analogously.\n\nAlternatively (to save space and compute ressources) one could try to only use the BERT_emb_[CLS] as these should contain information about the entire caption.\n\n**Step 3 (Train model):**\nTrain a MLP on this database.\n\nAlternatively (to save compute ressources) one could build a model f that maps the BERT embeddings to R^2048, and then one simply computes the sigmoid of the scalar product (Img_emb,f(BERT_emb)).  \n\n**Step 4 (Inference):**\nWhen we want to find the caption for some test image, we take the embedding of the image and match it to each test caption (e.g. compute \n1/N * (MLP(Img_emb,BERT_emb_word1) + ... + MLP(Img_emb,BERT_emb_wordN)) ) and predict the captions that match best.\n\nWhat do you think of this strategy?\n\nBest\nTobias\n\n___________________\n\nCode to get BERT embeddings:\n```\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport datatable as dt\nimport torch \nfrom transformers import BertTokenizer, BertModel\n\ndf = dt.fread('../input/wikipedia-image-caption/train-00000-of-00005.tsv')\n\ntokenizer = BertTokenizer.from_pretrained('bert-base-multilingual-uncased')\nmodel = BertModel.from_pretrained(\"bert-base-multilingual-uncased\", output_hidden_states = True)\nmodel.eval()\n\nntest = 10\ncapts = [capt.replace(' [SEP]','').replace('\\n',' ') for capt in df[:ntest,'caption_title_and_reference_description'].to_list()[0]]\nwith torch.no_grad():\n    for capt in capts:\n        print(capt)\n        tokenized = tokenizer.tokenize(\"[CLS] \" + capt + \" [SEP]\")\n        print(tokenized)\n        indexed = tokenizer.convert_tokens_to_ids(tokenized)\n        print(indexed)\n        segments = [1] * len(tokenized) #as a single sentence is fed to Bert, just put 1 for each token\n    \n        tokens = torch.tensor([indexed])\n        segments = torch.tensor([segments])\n        \n        outputs = model(tokens, segments)\n        hidden_states = outputs[2]\n        # tuple to tensor, get rid of batch dim, permute to #tokens x 13 hidden states x 768 model dim\n        token_embeddings = torch.stack(hidden_states, dim=0).squeeze(dim=1).permute(1,0,2)\n        # concatenate last 4 hiden states\n        token_embeddings = torch.flatten(token_embeddings[:,-4:],start_dim=1)\n        print(token_embeddings.size())\n```",
      "votes": null
    },
    {
      "id": "1566464",
      "postDate": "10/31/2021 22:20:46",
      "content": "<p>Hey I had thought of a similar approach. I wrote the code with the entire pipeline but turns out I can't download the images properly lol</p>",
      "rawMarkdown": "Hey I had thought of a similar approach. I wrote the code with the entire pipeline but turns out I can't download the images properly lol",
      "votes": null
    },
    {
      "id": "1568850",
      "postDate": "11/03/2021 04:43:45",
      "content": "<p>I am now going for a similar approach. The problem is that the data for image embedding has different order with the training .tsv files(that contains the captions). This makes it pretty hard to construct a database as you wrote above, since we have to make a match between 5 .tsv files with captions and 215 .csv files with image embeddings. I tried to just match them brute-force, but it took so much time to even make a few thousand rows of data :( I wonder if there would be a better way to handle this.</p>",
      "rawMarkdown": "I am now going for a similar approach. The problem is that the data for image embedding has different order with the training .tsv files(that contains the captions). This makes it pretty hard to construct a database as you wrote above, since we have to make a match between 5 .tsv files with captions and 215 .csv files with image embeddings. I tried to just match them brute-force, but it took so much time to even make a few thousand rows of data :( I wonder if there would be a better way to handle this.",
      "votes": null
    },
    {
      "id": "1568944",
      "postDate": "11/03/2021 06:46:45",
      "content": "<p>You don't need to match the CSVs with the TSVs. The CSVs are complete in themselves<br>\nI have created a dataset with negative sampling as mentioned by this discussion post<br>\nYou can check the discussion <a href=\"https://www.kaggle.com/c/wikipedia-image-caption/discussion/284720#1568842\" target=\"_blank\">here</a><br>\nThe notebook used to create the dataset can be found <a href=\"https://www.kaggle.com/debarshichanda/eda-versioning-easy-to-use-dataset\" target=\"_blank\">here</a><br>\nAt present I only use 3 CSV files</p>",
      "rawMarkdown": "You don't need to match the CSVs with the TSVs. The CSVs are complete in themselves\nI have created a dataset with negative sampling as mentioned by this discussion post\nYou can check the discussion [here](https://www.kaggle.com/c/wikipedia-image-caption/discussion/284720#1568842)\nThe notebook used to create the dataset can be found [here](https://www.kaggle.com/debarshichanda/eda-versioning-easy-to-use-dataset)\nAt present I only use 3 CSV files",
      "votes": null
    },
    {
      "id": "1568980",
      "postDate": "11/03/2021 07:09:33",
      "content": "<p>Wow. I was thinking of using only the resnet embeddings so I only downloaded the files in '/resnet_embeddings' directory. They didn't have any information about the caption…  guess I should've checked the joined data too. Thank you so much for contributing to the community!</p>",
      "rawMarkdown": "Wow. I was thinking of using only the resnet embeddings so I only downloaded the files in '/resnet_embeddings' directory. They didn't have any information about the caption...  guess I should've checked the joined data too. Thank you so much for contributing to the community!",
      "votes": null
    },
    {
      "id": "1570195",
      "postDate": "11/04/2021 01:19:09",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1566464,
      "author_name": "debarshichanda",
      "author_url": "",
      "post_date": "10/31/2021 22:20:46",
      "content": "<p>Hey I had thought of a similar approach. I wrote the code with the entire pipeline but turns out I can't download the images properly lol</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1568850,
      "author_name": "jiookchung",
      "author_url": "",
      "post_date": "11/03/2021 04:43:45",
      "content": "<p>I am now going for a similar approach. The problem is that the data for image embedding has different order with the training .tsv files(that contains the captions). This makes it pretty hard to construct a database as you wrote above, since we have to make a match between 5 .tsv files with captions and 215 .csv files with image embeddings. I tried to just match them brute-force, but it took so much time to even make a few thousand rows of data :( I wonder if there would be a better way to handle this.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1568944,
          "author_name": "debarshichanda",
          "author_url": "",
          "post_date": "11/03/2021 06:46:45",
          "content": "<p>You don't need to match the CSVs with the TSVs. The CSVs are complete in themselves<br>\nI have created a dataset with negative sampling as mentioned by this discussion post<br>\nYou can check the discussion <a href=\"https://www.kaggle.com/c/wikipedia-image-caption/discussion/284720#1568842\" target=\"_blank\">here</a><br>\nThe notebook used to create the dataset can be found <a href=\"https://www.kaggle.com/debarshichanda/eda-versioning-easy-to-use-dataset\" target=\"_blank\">here</a><br>\nAt present I only use 3 CSV files</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1568980,
          "author_name": "jiookchung",
          "author_url": "",
          "post_date": "11/03/2021 07:09:33",
          "content": "<p>Wow. I was thinking of using only the resnet embeddings so I only downloaded the files in '/resnet_embeddings' directory. They didn't have any information about the caption…  guess I should've checked the joined data too. Thank you so much for contributing to the community!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1570195,
          "author_name": "debarshichanda",
          "author_url": "",
          "post_date": "11/04/2021 01:19:09",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1540223": "Hi everyone!\n\nI wanted to share a strategy that I would have followed if I currently had more time and the necessary compute resources. I'd be happy to hear your thoughts about the strategy (maybe it's even the completely standard approach?) and maybe someone with more time and more compute resources wants to try it and can report his results.\n\nSo here is my idea:\n**Step 1 (Word embeddings):**\nUse the bert-base-multilingual-uncased model to obtain embeddings for the captions (e.g. \"Pariser Kanonen\" is tokenized into ['[CLS]', 'pariser', 'kanon', '##en', '[SEP]'] and the BERT model then yields for each of these 5 tokens an embedding, e.g. of dimension 768).\n\nFor more details how this works, I recommend the following blog post:\n[https://mccormickml.com/2019/05/14/BERT-word-embeddings-tutorial/](url)\nFurther, you may use the code provided in the end of this post.\n\n**Step 2 (Create Database):**\nCreate a database like the following:\n| Input 1 | Input 2 | Target |\n| --- | --- | --- |\n| Img_emb_PariserKanon  | BERT_emb_[CLS] | 1 |\n| Img_emb_PariserKanon  | BERT_emb_pariser | 1 |\n| Img_emb_PariserKanon  | BERT_emb_kanon | 1 |\n| Img_emb_PariserKanon  | BERT_emb_##en | 1 |\n| Img_emb_PariserKanon  | BERT_emb_[SEP] | 1 |\n| Img_emb_PariserKanon  | BERT_emb_[word from some other caption]_1 | 0 |\n| ... | ... | ... |\n| Img_emb_PariserKanon  | BERT_emb_[word from some other caption]_kk | 0 |\n\nWhere kk is some parameter, e.g. 20, for \"negative sampling\" (compare to word2Vec ideas, see e.g. this great blog post [https://jalammar.github.io/illustrated-word2vec/](url) ). And the embeddings of all other images and captions are inserted analogously.\n\nAlternatively (to save space and compute ressources) one could try to only use the BERT_emb_[CLS] as these should contain information about the entire caption.\n\n**Step 3 (Train model):**\nTrain a MLP on this database.\n\nAlternatively (to save compute ressources) one could build a model f that maps the BERT embeddings to R^2048, and then one simply computes the sigmoid of the scalar product (Img_emb,f(BERT_emb)).  \n\n**Step 4 (Inference):**\nWhen we want to find the caption for some test image, we take the embedding of the image and match it to each test caption (e.g. compute \n1/N * (MLP(Img_emb,BERT_emb_word1) + ... + MLP(Img_emb,BERT_emb_wordN)) ) and predict the captions that match best.\n\nWhat do you think of this strategy?\n\nBest\nTobias\n\n___________________\n\nCode to get BERT embeddings:\n```\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport datatable as dt\nimport torch \nfrom transformers import BertTokenizer, BertModel\n\ndf = dt.fread('../input/wikipedia-image-caption/train-00000-of-00005.tsv')\n\ntokenizer = BertTokenizer.from_pretrained('bert-base-multilingual-uncased')\nmodel = BertModel.from_pretrained(\"bert-base-multilingual-uncased\", output_hidden_states = True)\nmodel.eval()\n\nntest = 10\ncapts = [capt.replace(' [SEP]','').replace('\\n',' ') for capt in df[:ntest,'caption_title_and_reference_description'].to_list()[0]]\nwith torch.no_grad():\n    for capt in capts:\n        print(capt)\n        tokenized = tokenizer.tokenize(\"[CLS] \" + capt + \" [SEP]\")\n        print(tokenized)\n        indexed = tokenizer.convert_tokens_to_ids(tokenized)\n        print(indexed)\n        segments = [1] * len(tokenized) #as a single sentence is fed to Bert, just put 1 for each token\n    \n        tokens = torch.tensor([indexed])\n        segments = torch.tensor([segments])\n        \n        outputs = model(tokens, segments)\n        hidden_states = outputs[2]\n        # tuple to tensor, get rid of batch dim, permute to #tokens x 13 hidden states x 768 model dim\n        token_embeddings = torch.stack(hidden_states, dim=0).squeeze(dim=1).permute(1,0,2)\n        # concatenate last 4 hiden states\n        token_embeddings = torch.flatten(token_embeddings[:,-4:],start_dim=1)\n        print(token_embeddings.size())\n```",
    "1566464": "Hey I had thought of a similar approach. I wrote the code with the entire pipeline but turns out I can't download the images properly lol",
    "1568850": "I am now going for a similar approach. The problem is that the data for image embedding has different order with the training .tsv files(that contains the captions). This makes it pretty hard to construct a database as you wrote above, since we have to make a match between 5 .tsv files with captions and 215 .csv files with image embeddings. I tried to just match them brute-force, but it took so much time to even make a few thousand rows of data :( I wonder if there would be a better way to handle this.",
    "1568944": "You don't need to match the CSVs with the TSVs. The CSVs are complete in themselves\nI have created a dataset with negative sampling as mentioned by this discussion post\nYou can check the discussion [here](https://www.kaggle.com/c/wikipedia-image-caption/discussion/284720#1568842)\nThe notebook used to create the dataset can be found [here](https://www.kaggle.com/debarshichanda/eda-versioning-easy-to-use-dataset)\nAt present I only use 3 CSV files",
    "1568980": "Wow. I was thinking of using only the resnet embeddings so I only downloaded the files in '/resnet_embeddings' directory. They didn't have any information about the caption...  guess I should've checked the joined data too. Thank you so much for contributing to the community!",
    "1570195": "Thank you!"
  },
  "source": "meta"
}