{
  "id": 227753,
  "title": "Using ResNet for encoding and Transformer for decoding",
  "url": "/competitions/bms-molecular-translation/discussion/227753",
  "author_name": "",
  "post_date": "2021-03-22T05:07:50.358087500Z",
  "votes": 10,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hey everyone. I was thinking if it is possible to encode the image normally using a ResNet (or something similar) but then feed this representation to a Transformer for decoding. For example, you get a (N, 512, 7, 7) tensor out of the ResNet and then treat it like it is embeddings for 49 (7 * 7) words and feed it to the Transformer to do a kind of translation.<br>\nIn theory it seems plausible for me but I don't know if there could be any reason to do this.<br>\nI'll be really glad if you let me know what you think.</p>",
  "messages": [
    {
      "id": "1247826",
      "postDate": "03/22/2021 05:07:50",
      "content": "<p>Hey everyone. I was thinking if it is possible to encode the image normally using a ResNet (or something similar) but then feed this representation to a Transformer for decoding. For example, you get a (N, 512, 7, 7) tensor out of the ResNet and then treat it like it is embeddings for 49 (7 * 7) words and feed it to the Transformer to do a kind of translation.<br>\nIn theory it seems plausible for me but I don't know if there could be any reason to do this.<br>\nI'll be really glad if you let me know what you think.</p>",
      "rawMarkdown": "Hey everyone. I was thinking if it is possible to encode the image normally using a ResNet (or something similar) but then feed this representation to a Transformer for decoding. For example, you get a (N, 512, 7, 7) tensor out of the ResNet and then treat it like it is embeddings for 49 (7 * 7) words and feed it to the Transformer to do a kind of translation.\nIn theory it seems plausible for me but I don't know if there could be any reason to do this.\nI'll be really glad if you let me know what you think.",
      "votes": null
    },
    {
      "id": "1249918",
      "postDate": "03/23/2021 16:15:37",
      "content": "<p>Of course you can. </p>\n<p>The dataset is huge. LSTM is slow to train if you don't have descent hardware like myself. </p>",
      "rawMarkdown": "Of course you can. \n\nThe dataset is huge. LSTM is slow to train if you don't have descent hardware like myself.",
      "votes": null
    },
    {
      "id": "1249942",
      "postDate": "03/23/2021 16:35:30",
      "content": "<p>Exactly. My main problem with LSTM is that it is tooo slow and makes me unable to run many experiments on my available hardware. Good to hear that. I'm gonna start working on it.</p>",
      "rawMarkdown": "Exactly. My main problem with LSTM is that it is tooo slow and makes me unable to run many experiments on my available hardware. Good to hear that. I'm gonna start working on it.",
      "votes": null
    },
    {
      "id": "1250064",
      "postDate": "03/23/2021 18:42:01",
      "content": "<p>Especially if you plan to use TPU, I suspect it would be more efficient with a transformer as it is not sequential as LSTM/GRU</p>",
      "rawMarkdown": "Especially if you plan to use TPU, I suspect it would be more efficient with a transformer as it is not sequential as LSTM/GRU",
      "votes": null
    },
    {
      "id": "1250186",
      "postDate": "03/23/2021 20:25:52",
      "content": "<p>Oh, that's a good point. I did not know about it. Thanks for mentioning!</p>",
      "rawMarkdown": "Oh, that's a good point. I did not know about it. Thanks for mentioning!",
      "votes": null
    },
    {
      "id": "1250238",
      "postDate": "03/23/2021 21:20:39",
      "content": "<p>yes.. you just have to make sure to add  <code>positional embedding</code> to image features.. otherwise results will be funny =)</p>",
      "rawMarkdown": "yes.. you just have to make sure to add  `positional embedding` to image features.. otherwise results will be funny =)",
      "votes": null
    },
    {
      "id": "1250479",
      "postDate": "03/24/2021 04:37:06",
      "content": "<p>yes, sure. Do you think 1D positional embedding suffices or I must add 2D positional embedding? By 1D I mean I first flatten the 7 by 7 output grid and then treat them like sequential words in a sentence.</p>",
      "rawMarkdown": "yes, sure. Do you think 1D positional embedding suffices or I must add 2D positional embedding? By 1D I mean I first flatten the 7 by 7 output grid and then treat them like sequential words in a sentence.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1249918,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "03/23/2021 16:15:37",
      "content": "<p>Of course you can. </p>\n<p>The dataset is huge. LSTM is slow to train if you don't have descent hardware like myself. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1249942,
          "author_name": "moeinshariatnia",
          "author_url": "",
          "post_date": "03/23/2021 16:35:30",
          "content": "<p>Exactly. My main problem with LSTM is that it is tooo slow and makes me unable to run many experiments on my available hardware. Good to hear that. I'm gonna start working on it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1250064,
          "author_name": "ludovick",
          "author_url": "",
          "post_date": "03/23/2021 18:42:01",
          "content": "<p>Especially if you plan to use TPU, I suspect it would be more efficient with a transformer as it is not sequential as LSTM/GRU</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1250186,
          "author_name": "moeinshariatnia",
          "author_url": "",
          "post_date": "03/23/2021 20:25:52",
          "content": "<p>Oh, that's a good point. I did not know about it. Thanks for mentioning!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1250238,
      "author_name": "drhabib",
      "author_url": "",
      "post_date": "03/23/2021 21:20:39",
      "content": "<p>yes.. you just have to make sure to add  <code>positional embedding</code> to image features.. otherwise results will be funny =)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1250479,
          "author_name": "moeinshariatnia",
          "author_url": "",
          "post_date": "03/24/2021 04:37:06",
          "content": "<p>yes, sure. Do you think 1D positional embedding suffices or I must add 2D positional embedding? By 1D I mean I first flatten the 7 by 7 output grid and then treat them like sequential words in a sentence.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1247826": "Hey everyone. I was thinking if it is possible to encode the image normally using a ResNet (or something similar) but then feed this representation to a Transformer for decoding. For example, you get a (N, 512, 7, 7) tensor out of the ResNet and then treat it like it is embeddings for 49 (7 * 7) words and feed it to the Transformer to do a kind of translation.\nIn theory it seems plausible for me but I don't know if there could be any reason to do this.\nI'll be really glad if you let me know what you think.",
    "1249918": "Of course you can. \n\nThe dataset is huge. LSTM is slow to train if you don't have descent hardware like myself.",
    "1249942": "Exactly. My main problem with LSTM is that it is tooo slow and makes me unable to run many experiments on my available hardware. Good to hear that. I'm gonna start working on it.",
    "1250064": "Especially if you plan to use TPU, I suspect it would be more efficient with a transformer as it is not sequential as LSTM/GRU",
    "1250186": "Oh, that's a good point. I did not know about it. Thanks for mentioning!",
    "1250238": "yes.. you just have to make sure to add  `positional embedding` to image features.. otherwise results will be funny =)",
    "1250479": "yes, sure. Do you think 1D positional embedding suffices or I must add 2D positional embedding? By 1D I mean I first flatten the 7 by 7 output grid and then treat them like sequential words in a sentence."
  },
  "source": "meta"
}