{
  "id": 223212,
  "title": "End-to-end transformers approach with ViT and vanilla transformer decoder (idea)",
  "url": "/competitions/bms-molecular-translation/discussion/223212",
  "author_name": "",
  "post_date": "2021-03-03T00:22:05.735798200Z",
  "votes": 24,
  "comment_count": 18,
  "views": 0,
  "content": "<p>This post does not describe a working system.</p>\n<p>I think it would be interesting to see how well an end-to-end transformer approach would do here. As you know, ViT is capable of achieving excellent performance by taking patches of an image and pass them through a transformer encoder for classification. If you take it one step further, you could use a pre-trained ViT model as the encoder (i.e. use the output feature maps instead of the classification head) and pass it through a transformer decoder (which could be using the vanilla transformer architecture). Then, you could auto-regressively generate one character at the time (or perhaps you could tokenize each label to have meaningful sub-sequences to be decoded).</p>\n<p>For the decoder part, you might not be able to leverage pre-training, unless you use something like BART's decoder, but there are no guarantees it would work. However, an end-to-end transformer architecture could potentially make code simpler and less hyper-parameters to tune. </p>\n<p>Here's a sketch of what i have in mind (taken from the BART and ViT papers):</p>\n<p><img src=\"https://i.imgur.com/s6tnVnT.png\" alt=\"\"></p>",
  "messages": [
    {
      "id": "1224657",
      "postDate": "03/03/2021 00:22:05",
      "content": "<p>This post does not describe a working system.</p>\n<p>I think it would be interesting to see how well an end-to-end transformer approach would do here. As you know, ViT is capable of achieving excellent performance by taking patches of an image and pass them through a transformer encoder for classification. If you take it one step further, you could use a pre-trained ViT model as the encoder (i.e. use the output feature maps instead of the classification head) and pass it through a transformer decoder (which could be using the vanilla transformer architecture). Then, you could auto-regressively generate one character at the time (or perhaps you could tokenize each label to have meaningful sub-sequences to be decoded).</p>\n<p>For the decoder part, you might not be able to leverage pre-training, unless you use something like BART's decoder, but there are no guarantees it would work. However, an end-to-end transformer architecture could potentially make code simpler and less hyper-parameters to tune. </p>\n<p>Here's a sketch of what i have in mind (taken from the BART and ViT papers):</p>\n<p><img src=\"https://i.imgur.com/s6tnVnT.png\" alt=\"\"></p>",
      "rawMarkdown": "This post does not describe a working system.\n\nI think it would be interesting to see how well an end-to-end transformer approach would do here. As you know, ViT is capable of achieving excellent performance by taking patches of an image and pass them through a transformer encoder for classification. If you take it one step further, you could use a pre-trained ViT model as the encoder (i.e. use the output feature maps instead of the classification head) and pass it through a transformer decoder (which could be using the vanilla transformer architecture). Then, you could auto-regressively generate one character at the time (or perhaps you could tokenize each label to have meaningful sub-sequences to be decoded).\n\nFor the decoder part, you might not be able to leverage pre-training, unless you use something like BART's decoder, but there are no guarantees it would work. However, an end-to-end transformer architecture could potentially make code simpler and less hyper-parameters to tune. \n\nHere's a sketch of what i have in mind (taken from the BART and ViT papers):\n\n![](https://i.imgur.com/s6tnVnT.png)",
      "votes": null
    },
    {
      "id": "1224681",
      "postDate": "03/03/2021 01:21:28",
      "content": "<p>this approach is correct, but maybe we cannot simply divide the image by patch </p>\n<p>some smarter algorithm needs to cut the image correctly.</p>",
      "rawMarkdown": "this approach is correct, but maybe we cannot simply divide the image by patch \n\n some smarter algorithm needs to cut the image correctly.",
      "votes": null
    },
    {
      "id": "1224686",
      "postDate": "03/03/2021 01:27:22",
      "content": "<p>interesting. i wonder if that segmentation could be learned in a unsupervised manner, or if some sort of heuristic might work better in this case.</p>",
      "rawMarkdown": "interesting. i wonder if that segmentation could be learned in a unsupervised manner, or if some sort of heuristic might work better in this case.",
      "votes": null
    },
    {
      "id": "1224687",
      "postDate": "03/03/2021 01:27:30",
      "content": "<p>The patch should be multi-scaled, just like WLET</p>",
      "rawMarkdown": "The patch should be multi-scaled, just like WLET",
      "votes": null
    },
    {
      "id": "1224886",
      "postDate": "03/03/2021 06:30:52",
      "content": "<p>\" segmentation could be learned in a unsupervised manner,\"</p>\n<p>a smarter way is to use python tools (e.g. molecular/chemistry lib) to convert from text to image (diargram). such tools allows us to output the image format. e.g. we can output ring, functional group in a certain color, etc these color becomes labels for learning segmentation</p>\n<p>this is a competition where the shakeup can be controlled.</p>\n<p>the first that comes to my mind is how to come up with a script to generate all the possible(or most of it) molecular formulas? I think the number is large but not infinite</p>",
      "rawMarkdown": "\" segmentation could be learned in a unsupervised manner,\"\n\na smarter way is to use python tools (e.g. molecular/chemistry lib) to convert from text to image (diargram). such tools allows us to output the image format. e.g. we can output ring, functional group in a certain color, etc these color becomes labels for learning segmentation\n\nthis is a competition where the shakeup can be controlled.\n\nthe first that comes to my mind is how to come up with a script to generate all the possible(or most of it) molecular formulas? I think the number is large but not infinite",
      "votes": null
    },
    {
      "id": "1224976",
      "postDate": "03/03/2021 08:04:10",
      "content": "<p>I'm not actually sure a multi-scale architecture is needed for this problem. Seems to me that all the images are on the same relative scale.</p>",
      "rawMarkdown": "I'm not actually sure a multi-scale architecture is needed for this problem. Seems to me that all the images are on the same relative scale.",
      "votes": null
    },
    {
      "id": "1225768",
      "postDate": "03/03/2021 22:56:41",
      "content": "<p>it is used to solve the patch size problem </p>",
      "rawMarkdown": "it is used to solve the patch size problem",
      "votes": null
    },
    {
      "id": "1225797",
      "postDate": "03/04/2021 00:07:48",
      "content": "<p>Love the Hinton-style opening :-)</p>",
      "rawMarkdown": "Love the Hinton-style opening :-)",
      "votes": null
    },
    {
      "id": "1225868",
      "postDate": "03/04/2021 02:49:57",
      "content": "<p>Only Hinton can be trendy after just a few days on arxiv :))</p>",
      "rawMarkdown": "Only Hinton can be trendy after just a few days on arxiv :))",
      "votes": null
    },
    {
      "id": "1225969",
      "postDate": "03/04/2021 05:32:13",
      "content": "<p>I was wrong any ways. It appears some images have been scaled down ~1/2</p>",
      "rawMarkdown": "I was wrong any ways. It appears some images have been scaled down ~1/2",
      "votes": null
    },
    {
      "id": "1226022",
      "postDate": "03/04/2021 06:57:45",
      "content": "<p>there is a dense transformer version:<br>\n<a href=\"https://arxiv.org/abs/2102.12122\" target=\"_blank\">https://arxiv.org/abs/2102.12122</a><br>\nPyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions</p>",
      "rawMarkdown": "there is a dense transformer version:\nhttps://arxiv.org/abs/2102.12122\nPyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions",
      "votes": null
    },
    {
      "id": "1227566",
      "postDate": "03/05/2021 16:47:38",
      "content": "<p>That's pretty interesting, but how do we ensure that the synthetically generated images will be similar enough to the source images in this competition in order to zero-shot transfer the learned segmentations?</p>",
      "rawMarkdown": "That's pretty interesting, but how do we ensure that the synthetically generated images will be similar enough to the source images in this competition in order to zero-shot transfer the learned segmentations?",
      "votes": null
    },
    {
      "id": "1227918",
      "postDate": "03/05/2021 23:56:52",
      "content": "<p>Interesting to note this <a href=\"https://www.kaggle.com/tj0612/convert-from-inchi-to-picture-by-rdkit\" target=\"_blank\">excellent notebook</a> showing how to convert InChI to images with rdkit.</p>",
      "rawMarkdown": "Interesting to note this [excellent notebook](https://www.kaggle.com/tj0612/convert-from-inchi-to-picture-by-rdkit) showing how to convert InChI to images with rdkit.",
      "votes": null
    },
    {
      "id": "1228027",
      "postDate": "03/06/2021 03:48:38",
      "content": "<p>As far as I understand, this conversion does work but the orientation of the source image might be different from the generated one something like this: </p>\n<p><img src=\"https://www.kaggleusercontent.com/kf/55817187/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..CLP1N_ZNjoswCtzS0amNrA.I7klGdfBaOZUPv_0PALYp1WhSuJ1HFDD0Ks43NXcDQZ33rvR-lqErIDS6asMkcNMuKQaoBr7UyUlAxvn2yWPrGemNP0zY5DtZdBSxP3XyfrIRXTAzNTDu1s-B9Au7zS91y92n0z2c6Ivr_yFBHhWMzHMVZwLMovUfcMOWhBRkNkmuhNJ2o4sQfs5PmUospdKluTGHpfx4SRwJd-yfgxxUYR0PI_nwteioihbfFcSHNqa6Qq1PJEgZ20pQuZ2XVPRzts0jDLKTFynCQVnw_0lk_Vifx6Kr3Uz_tQ95QJ7_4DOQbIRN2ViWuo_lYSO-5wdWJGVWfm1oQagbt-fV1FunuBhInr88DWBQ63bpBFkpVafZ5GxboiZMYqu_w9tbrkEm2VsLE5V-2uwEXtG8tV75wzIm0UsQsc1pvB1aQP2AI4MUpF0u47GIAWeCvuZdXcQ4Wlbhw-eFrvbhA_x-qjyN3rfsbQWA9ipMVUZpYhXvuIcYaDL3OG1d31JC4eJNI9biOdzXm_fHPkIk6Tl4Zt-r_qOJI8CrbF6KQktvogN7zmyLcF2iN0vWe9ijDLDvmBUjJVZs7yAvKAa-NBYwmmRtFTlpCbf2tLjJyUqRPnzs6sPkWe8BBFadpHbu7TeQBRygpjbXeMeYNB5yjholWclWF26157w9N01GCLomapxRu5FvvSFWjYkrSphcuUpE4Cm.cE5gttNDWTZ2O1vmmP6QzQ/__results___files/__results___6_1.png\" alt=\"\"></p>\n<p><img src=\"https://www.kaggleusercontent.com/kf/55817187/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..CLP1N_ZNjoswCtzS0amNrA.I7klGdfBaOZUPv_0PALYp1WhSuJ1HFDD0Ks43NXcDQZ33rvR-lqErIDS6asMkcNMuKQaoBr7UyUlAxvn2yWPrGemNP0zY5DtZdBSxP3XyfrIRXTAzNTDu1s-B9Au7zS91y92n0z2c6Ivr_yFBHhWMzHMVZwLMovUfcMOWhBRkNkmuhNJ2o4sQfs5PmUospdKluTGHpfx4SRwJd-yfgxxUYR0PI_nwteioihbfFcSHNqa6Qq1PJEgZ20pQuZ2XVPRzts0jDLKTFynCQVnw_0lk_Vifx6Kr3Uz_tQ95QJ7_4DOQbIRN2ViWuo_lYSO-5wdWJGVWfm1oQagbt-fV1FunuBhInr88DWBQ63bpBFkpVafZ5GxboiZMYqu_w9tbrkEm2VsLE5V-2uwEXtG8tV75wzIm0UsQsc1pvB1aQP2AI4MUpF0u47GIAWeCvuZdXcQ4Wlbhw-eFrvbhA_x-qjyN3rfsbQWA9ipMVUZpYhXvuIcYaDL3OG1d31JC4eJNI9biOdzXm_fHPkIk6Tl4Zt-r_qOJI8CrbF6KQktvogN7zmyLcF2iN0vWe9ijDLDvmBUjJVZs7yAvKAa-NBYwmmRtFTlpCbf2tLjJyUqRPnzs6sPkWe8BBFadpHbu7TeQBRygpjbXeMeYNB5yjholWclWF26157w9N01GCLomapxRu5FvvSFWjYkrSphcuUpE4Cm.cE5gttNDWTZ2O1vmmP6QzQ/__results___files/__results___6_2.png\" alt=\"\"></p>\n<p>How exactly can we learn segmentation from differently oriented source and mask? </p>",
      "rawMarkdown": "As far as I understand, this conversion does work but the orientation of the source image might be different from the generated one something like this: \n\n\n\n\n![](https://www.kaggleusercontent.com/kf/55817187/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..CLP1N_ZNjoswCtzS0amNrA.I7klGdfBaOZUPv_0PALYp1WhSuJ1HFDD0Ks43NXcDQZ33rvR-lqErIDS6asMkcNMuKQaoBr7UyUlAxvn2yWPrGemNP0zY5DtZdBSxP3XyfrIRXTAzNTDu1s-B9Au7zS91y92n0z2c6Ivr_yFBHhWMzHMVZwLMovUfcMOWhBRkNkmuhNJ2o4sQfs5PmUospdKluTGHpfx4SRwJd-yfgxxUYR0PI_nwteioihbfFcSHNqa6Qq1PJEgZ20pQuZ2XVPRzts0jDLKTFynCQVnw_0lk_Vifx6Kr3Uz_tQ95QJ7_4DOQbIRN2ViWuo_lYSO-5wdWJGVWfm1oQagbt-fV1FunuBhInr88DWBQ63bpBFkpVafZ5GxboiZMYqu_w9tbrkEm2VsLE5V-2uwEXtG8tV75wzIm0UsQsc1pvB1aQP2AI4MUpF0u47GIAWeCvuZdXcQ4Wlbhw-eFrvbhA_x-qjyN3rfsbQWA9ipMVUZpYhXvuIcYaDL3OG1d31JC4eJNI9biOdzXm_fHPkIk6Tl4Zt-r_qOJI8CrbF6KQktvogN7zmyLcF2iN0vWe9ijDLDvmBUjJVZs7yAvKAa-NBYwmmRtFTlpCbf2tLjJyUqRPnzs6sPkWe8BBFadpHbu7TeQBRygpjbXeMeYNB5yjholWclWF26157w9N01GCLomapxRu5FvvSFWjYkrSphcuUpE4Cm.cE5gttNDWTZ2O1vmmP6QzQ/__results___files/__results___6_1.png)\n\n![](https://www.kaggleusercontent.com/kf/55817187/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..CLP1N_ZNjoswCtzS0amNrA.I7klGdfBaOZUPv_0PALYp1WhSuJ1HFDD0Ks43NXcDQZ33rvR-lqErIDS6asMkcNMuKQaoBr7UyUlAxvn2yWPrGemNP0zY5DtZdBSxP3XyfrIRXTAzNTDu1s-B9Au7zS91y92n0z2c6Ivr_yFBHhWMzHMVZwLMovUfcMOWhBRkNkmuhNJ2o4sQfs5PmUospdKluTGHpfx4SRwJd-yfgxxUYR0PI_nwteioihbfFcSHNqa6Qq1PJEgZ20pQuZ2XVPRzts0jDLKTFynCQVnw_0lk_Vifx6Kr3Uz_tQ95QJ7_4DOQbIRN2ViWuo_lYSO-5wdWJGVWfm1oQagbt-fV1FunuBhInr88DWBQ63bpBFkpVafZ5GxboiZMYqu_w9tbrkEm2VsLE5V-2uwEXtG8tV75wzIm0UsQsc1pvB1aQP2AI4MUpF0u47GIAWeCvuZdXcQ4Wlbhw-eFrvbhA_x-qjyN3rfsbQWA9ipMVUZpYhXvuIcYaDL3OG1d31JC4eJNI9biOdzXm_fHPkIk6Tl4Zt-r_qOJI8CrbF6KQktvogN7zmyLcF2iN0vWe9ijDLDvmBUjJVZs7yAvKAa-NBYwmmRtFTlpCbf2tLjJyUqRPnzs6sPkWe8BBFadpHbu7TeQBRygpjbXeMeYNB5yjholWclWF26157w9N01GCLomapxRu5FvvSFWjYkrSphcuUpE4Cm.cE5gttNDWTZ2O1vmmP6QzQ/__results___files/__results___6_2.png)\n\n\nHow exactly can we learn segmentation from differently oriented source and mask?",
      "votes": null
    },
    {
      "id": "1228057",
      "postDate": "03/06/2021 04:46:42",
      "content": "<p>Yes, of course it may be different. There is no one correct way to layout most molecules, and that is part of the challenge inherent to this competition. So you can't use the generated image as a mask directly, but perhaps there's other ways it can be of use ;)</p>",
      "rawMarkdown": "Yes, of course it may be different. There is no one correct way to layout most molecules, and that is part of the challenge inherent to this competition. So you can't use the generated image as a mask directly, but perhaps there's other ways it can be of use ;)",
      "votes": null
    },
    {
      "id": "1228649",
      "postDate": "03/06/2021 16:09:16",
      "content": "<p>Perhaps it's possible to segment the synthetic images based on the pre-drawing information (i.e. you will know ahead of time which what will be contained in each part of the image). then once you train a model that learns to segment those synthetic images, you can apply it directly on the images in the training/test set without training any further (i.e. zero-shot transfer learning).</p>",
      "rawMarkdown": "Perhaps it's possible to segment the synthetic images based on the pre-drawing information (i.e. you will know ahead of time which what will be contained in each part of the image). then once you train a model that learns to segment those synthetic images, you can apply it directly on the images in the training/test set without training any further (i.e. zero-shot transfer learning).",
      "votes": null
    },
    {
      "id": "1238103",
      "postDate": "03/14/2021 17:00:53",
      "content": "<p>I really like the idea of using only the Transformer model, however the ViT paper shows that, although the ViT model produces SOTA results on ImageNet, it requires much more data in order to beat traditional CNN models (eg. ResNet). Since the dataset is only approx. 4M images, unless you generate a very significant amount of additional data, it might be the case that traditional CNN models will work better.</p>",
      "rawMarkdown": "I really like the idea of using only the Transformer model, however the ViT paper shows that, although the ViT model produces SOTA results on ImageNet, it requires much more data in order to beat traditional CNN models (eg. ResNet). Since the dataset is only approx. 4M images, unless you generate a very significant amount of additional data, it might be the case that traditional CNN models will work better.",
      "votes": null
    },
    {
      "id": "1244412",
      "postDate": "03/19/2021 00:59:28",
      "content": "<p>Cassava competition had 27000 training images and <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/221150\" target=\"_blank\">the 3rd place solution</a> was an ensemble of ONLY ViT models! I think it's totally doable here with 4M images. <br>\nThe only limitation with ViT is the image resolution. If size matters in this competition then ViT won't be performant since the available pretrained models can't handle more than 384x384 image size.</p>",
      "rawMarkdown": "Cassava competition had 27000 training images and [the 3rd place solution](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/221150) was an ensemble of ONLY ViT models! I think it's totally doable here with 4M images. \nThe only limitation with ViT is the image resolution. If size matters in this competition then ViT won't be performant since the available pretrained models can't handle more than 384x384 image size.",
      "votes": null
    },
    {
      "id": "1260795",
      "postDate": "04/02/2021 12:17:19",
      "content": "<p>My model has this exact implementation (4.49 best LB score). More precisely I have a slightly modified vit_deit_tiny_distilled_patch16_224 from timm as encoder connected to 3 decoder layers (~30MB model size).</p>\n<p>I'm aiming now to increase my model size and improve my preprocessing and data augmentation pipeline.</p>\n<p>The bad thing is that my last model trained during ~75 epochs with 1h per epoch (RTX3090). For inference (autoregressive decoding), it needed around ~22h in float32 precision and ~13h in float16. So basically I can just make one submission a week. </p>",
      "rawMarkdown": "My model has this exact implementation (4.49 best LB score). More precisely I have a slightly modified vit_deit_tiny_distilled_patch16_224 from timm as encoder connected to 3 decoder layers (~30MB model size).\n\nI'm aiming now to increase my model size and improve my preprocessing and data augmentation pipeline.\n\nThe bad thing is that my last model trained during ~75 epochs with 1h per epoch (RTX3090). For inference (autoregressive decoding), it needed around ~22h in float32 precision and ~13h in float16. So basically I can just make one submission a week.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1224681,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/03/2021 01:21:28",
      "content": "<p>this approach is correct, but maybe we cannot simply divide the image by patch </p>\n<p>some smarter algorithm needs to cut the image correctly.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1224686,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "03/03/2021 01:27:22",
          "content": "<p>interesting. i wonder if that segmentation could be learned in a unsupervised manner, or if some sort of heuristic might work better in this case.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1224886,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/03/2021 06:30:52",
          "content": "<p>\" segmentation could be learned in a unsupervised manner,\"</p>\n<p>a smarter way is to use python tools (e.g. molecular/chemistry lib) to convert from text to image (diargram). such tools allows us to output the image format. e.g. we can output ring, functional group in a certain color, etc these color becomes labels for learning segmentation</p>\n<p>this is a competition where the shakeup can be controlled.</p>\n<p>the first that comes to my mind is how to come up with a script to generate all the possible(or most of it) molecular formulas? I think the number is large but not infinite</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1227566,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "03/05/2021 16:47:38",
          "content": "<p>That's pretty interesting, but how do we ensure that the synthetically generated images will be similar enough to the source images in this competition in order to zero-shot transfer the learned segmentations?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1227918,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "03/05/2021 23:56:52",
          "content": "<p>Interesting to note this <a href=\"https://www.kaggle.com/tj0612/convert-from-inchi-to-picture-by-rdkit\" target=\"_blank\">excellent notebook</a> showing how to convert InChI to images with rdkit.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1228027,
          "author_name": "arka47",
          "author_url": "",
          "post_date": "03/06/2021 03:48:38",
          "content": "<p>As far as I understand, this conversion does work but the orientation of the source image might be different from the generated one something like this: </p>\n<p><img src=\"https://www.kaggleusercontent.com/kf/55817187/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..CLP1N_ZNjoswCtzS0amNrA.I7klGdfBaOZUPv_0PALYp1WhSuJ1HFDD0Ks43NXcDQZ33rvR-lqErIDS6asMkcNMuKQaoBr7UyUlAxvn2yWPrGemNP0zY5DtZdBSxP3XyfrIRXTAzNTDu1s-B9Au7zS91y92n0z2c6Ivr_yFBHhWMzHMVZwLMovUfcMOWhBRkNkmuhNJ2o4sQfs5PmUospdKluTGHpfx4SRwJd-yfgxxUYR0PI_nwteioihbfFcSHNqa6Qq1PJEgZ20pQuZ2XVPRzts0jDLKTFynCQVnw_0lk_Vifx6Kr3Uz_tQ95QJ7_4DOQbIRN2ViWuo_lYSO-5wdWJGVWfm1oQagbt-fV1FunuBhInr88DWBQ63bpBFkpVafZ5GxboiZMYqu_w9tbrkEm2VsLE5V-2uwEXtG8tV75wzIm0UsQsc1pvB1aQP2AI4MUpF0u47GIAWeCvuZdXcQ4Wlbhw-eFrvbhA_x-qjyN3rfsbQWA9ipMVUZpYhXvuIcYaDL3OG1d31JC4eJNI9biOdzXm_fHPkIk6Tl4Zt-r_qOJI8CrbF6KQktvogN7zmyLcF2iN0vWe9ijDLDvmBUjJVZs7yAvKAa-NBYwmmRtFTlpCbf2tLjJyUqRPnzs6sPkWe8BBFadpHbu7TeQBRygpjbXeMeYNB5yjholWclWF26157w9N01GCLomapxRu5FvvSFWjYkrSphcuUpE4Cm.cE5gttNDWTZ2O1vmmP6QzQ/__results___files/__results___6_1.png\" alt=\"\"></p>\n<p><img src=\"https://www.kaggleusercontent.com/kf/55817187/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..CLP1N_ZNjoswCtzS0amNrA.I7klGdfBaOZUPv_0PALYp1WhSuJ1HFDD0Ks43NXcDQZ33rvR-lqErIDS6asMkcNMuKQaoBr7UyUlAxvn2yWPrGemNP0zY5DtZdBSxP3XyfrIRXTAzNTDu1s-B9Au7zS91y92n0z2c6Ivr_yFBHhWMzHMVZwLMovUfcMOWhBRkNkmuhNJ2o4sQfs5PmUospdKluTGHpfx4SRwJd-yfgxxUYR0PI_nwteioihbfFcSHNqa6Qq1PJEgZ20pQuZ2XVPRzts0jDLKTFynCQVnw_0lk_Vifx6Kr3Uz_tQ95QJ7_4DOQbIRN2ViWuo_lYSO-5wdWJGVWfm1oQagbt-fV1FunuBhInr88DWBQ63bpBFkpVafZ5GxboiZMYqu_w9tbrkEm2VsLE5V-2uwEXtG8tV75wzIm0UsQsc1pvB1aQP2AI4MUpF0u47GIAWeCvuZdXcQ4Wlbhw-eFrvbhA_x-qjyN3rfsbQWA9ipMVUZpYhXvuIcYaDL3OG1d31JC4eJNI9biOdzXm_fHPkIk6Tl4Zt-r_qOJI8CrbF6KQktvogN7zmyLcF2iN0vWe9ijDLDvmBUjJVZs7yAvKAa-NBYwmmRtFTlpCbf2tLjJyUqRPnzs6sPkWe8BBFadpHbu7TeQBRygpjbXeMeYNB5yjholWclWF26157w9N01GCLomapxRu5FvvSFWjYkrSphcuUpE4Cm.cE5gttNDWTZ2O1vmmP6QzQ/__results___files/__results___6_2.png\" alt=\"\"></p>\n<p>How exactly can we learn segmentation from differently oriented source and mask? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1228057,
          "author_name": "matthewmasters",
          "author_url": "",
          "post_date": "03/06/2021 04:46:42",
          "content": "<p>Yes, of course it may be different. There is no one correct way to layout most molecules, and that is part of the challenge inherent to this competition. So you can't use the generated image as a mask directly, but perhaps there's other ways it can be of use ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1228649,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "03/06/2021 16:09:16",
          "content": "<p>Perhaps it's possible to segment the synthetic images based on the pre-drawing information (i.e. you will know ahead of time which what will be contained in each part of the image). then once you train a model that learns to segment those synthetic images, you can apply it directly on the images in the training/test set without training any further (i.e. zero-shot transfer learning).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1224687,
      "author_name": "zzllusa",
      "author_url": "",
      "post_date": "03/03/2021 01:27:30",
      "content": "<p>The patch should be multi-scaled, just like WLET</p>",
      "votes": null,
      "replies": [
        {
          "id": 1224976,
          "author_name": "matthewmasters",
          "author_url": "",
          "post_date": "03/03/2021 08:04:10",
          "content": "<p>I'm not actually sure a multi-scale architecture is needed for this problem. Seems to me that all the images are on the same relative scale.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1225768,
          "author_name": "zzllusa",
          "author_url": "",
          "post_date": "03/03/2021 22:56:41",
          "content": "<p>it is used to solve the patch size problem </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1225969,
          "author_name": "matthewmasters",
          "author_url": "",
          "post_date": "03/04/2021 05:32:13",
          "content": "<p>I was wrong any ways. It appears some images have been scaled down ~1/2</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1226022,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/04/2021 06:57:45",
          "content": "<p>there is a dense transformer version:<br>\n<a href=\"https://arxiv.org/abs/2102.12122\" target=\"_blank\">https://arxiv.org/abs/2102.12122</a><br>\nPyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1225797,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "03/04/2021 00:07:48",
      "content": "<p>Love the Hinton-style opening :-)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1225868,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "03/04/2021 02:49:57",
          "content": "<p>Only Hinton can be trendy after just a few days on arxiv :))</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1238103,
      "author_name": "rssrwn",
      "author_url": "",
      "post_date": "03/14/2021 17:00:53",
      "content": "<p>I really like the idea of using only the Transformer model, however the ViT paper shows that, although the ViT model produces SOTA results on ImageNet, it requires much more data in order to beat traditional CNN models (eg. ResNet). Since the dataset is only approx. 4M images, unless you generate a very significant amount of additional data, it might be the case that traditional CNN models will work better.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1244412,
          "author_name": "amiiiney",
          "author_url": "",
          "post_date": "03/19/2021 00:59:28",
          "content": "<p>Cassava competition had 27000 training images and <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/221150\" target=\"_blank\">the 3rd place solution</a> was an ensemble of ONLY ViT models! I think it's totally doable here with 4M images. <br>\nThe only limitation with ViT is the image resolution. If size matters in this competition then ViT won't be performant since the available pretrained models can't handle more than 384x384 image size.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1260795,
      "author_name": "claverru",
      "author_url": "",
      "post_date": "04/02/2021 12:17:19",
      "content": "<p>My model has this exact implementation (4.49 best LB score). More precisely I have a slightly modified vit_deit_tiny_distilled_patch16_224 from timm as encoder connected to 3 decoder layers (~30MB model size).</p>\n<p>I'm aiming now to increase my model size and improve my preprocessing and data augmentation pipeline.</p>\n<p>The bad thing is that my last model trained during ~75 epochs with 1h per epoch (RTX3090). For inference (autoregressive decoding), it needed around ~22h in float32 precision and ~13h in float16. So basically I can just make one submission a week. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1224657": "This post does not describe a working system.\n\nI think it would be interesting to see how well an end-to-end transformer approach would do here. As you know, ViT is capable of achieving excellent performance by taking patches of an image and pass them through a transformer encoder for classification. If you take it one step further, you could use a pre-trained ViT model as the encoder (i.e. use the output feature maps instead of the classification head) and pass it through a transformer decoder (which could be using the vanilla transformer architecture). Then, you could auto-regressively generate one character at the time (or perhaps you could tokenize each label to have meaningful sub-sequences to be decoded).\n\nFor the decoder part, you might not be able to leverage pre-training, unless you use something like BART's decoder, but there are no guarantees it would work. However, an end-to-end transformer architecture could potentially make code simpler and less hyper-parameters to tune. \n\nHere's a sketch of what i have in mind (taken from the BART and ViT papers):\n\n![](https://i.imgur.com/s6tnVnT.png)",
    "1224681": "this approach is correct, but maybe we cannot simply divide the image by patch \n\n some smarter algorithm needs to cut the image correctly.",
    "1224686": "interesting. i wonder if that segmentation could be learned in a unsupervised manner, or if some sort of heuristic might work better in this case.",
    "1224687": "The patch should be multi-scaled, just like WLET",
    "1224886": "\" segmentation could be learned in a unsupervised manner,\"\n\na smarter way is to use python tools (e.g. molecular/chemistry lib) to convert from text to image (diargram). such tools allows us to output the image format. e.g. we can output ring, functional group in a certain color, etc these color becomes labels for learning segmentation\n\nthis is a competition where the shakeup can be controlled.\n\nthe first that comes to my mind is how to come up with a script to generate all the possible(or most of it) molecular formulas? I think the number is large but not infinite",
    "1224976": "I'm not actually sure a multi-scale architecture is needed for this problem. Seems to me that all the images are on the same relative scale.",
    "1225768": "it is used to solve the patch size problem",
    "1225797": "Love the Hinton-style opening :-)",
    "1225868": "Only Hinton can be trendy after just a few days on arxiv :))",
    "1225969": "I was wrong any ways. It appears some images have been scaled down ~1/2",
    "1226022": "there is a dense transformer version:\nhttps://arxiv.org/abs/2102.12122\nPyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions",
    "1227566": "That's pretty interesting, but how do we ensure that the synthetically generated images will be similar enough to the source images in this competition in order to zero-shot transfer the learned segmentations?",
    "1227918": "Interesting to note this [excellent notebook](https://www.kaggle.com/tj0612/convert-from-inchi-to-picture-by-rdkit) showing how to convert InChI to images with rdkit.",
    "1228027": "As far as I understand, this conversion does work but the orientation of the source image might be different from the generated one something like this: \n\n\n\n\n![](https://www.kaggleusercontent.com/kf/55817187/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..CLP1N_ZNjoswCtzS0amNrA.I7klGdfBaOZUPv_0PALYp1WhSuJ1HFDD0Ks43NXcDQZ33rvR-lqErIDS6asMkcNMuKQaoBr7UyUlAxvn2yWPrGemNP0zY5DtZdBSxP3XyfrIRXTAzNTDu1s-B9Au7zS91y92n0z2c6Ivr_yFBHhWMzHMVZwLMovUfcMOWhBRkNkmuhNJ2o4sQfs5PmUospdKluTGHpfx4SRwJd-yfgxxUYR0PI_nwteioihbfFcSHNqa6Qq1PJEgZ20pQuZ2XVPRzts0jDLKTFynCQVnw_0lk_Vifx6Kr3Uz_tQ95QJ7_4DOQbIRN2ViWuo_lYSO-5wdWJGVWfm1oQagbt-fV1FunuBhInr88DWBQ63bpBFkpVafZ5GxboiZMYqu_w9tbrkEm2VsLE5V-2uwEXtG8tV75wzIm0UsQsc1pvB1aQP2AI4MUpF0u47GIAWeCvuZdXcQ4Wlbhw-eFrvbhA_x-qjyN3rfsbQWA9ipMVUZpYhXvuIcYaDL3OG1d31JC4eJNI9biOdzXm_fHPkIk6Tl4Zt-r_qOJI8CrbF6KQktvogN7zmyLcF2iN0vWe9ijDLDvmBUjJVZs7yAvKAa-NBYwmmRtFTlpCbf2tLjJyUqRPnzs6sPkWe8BBFadpHbu7TeQBRygpjbXeMeYNB5yjholWclWF26157w9N01GCLomapxRu5FvvSFWjYkrSphcuUpE4Cm.cE5gttNDWTZ2O1vmmP6QzQ/__results___files/__results___6_1.png)\n\n![](https://www.kaggleusercontent.com/kf/55817187/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..CLP1N_ZNjoswCtzS0amNrA.I7klGdfBaOZUPv_0PALYp1WhSuJ1HFDD0Ks43NXcDQZ33rvR-lqErIDS6asMkcNMuKQaoBr7UyUlAxvn2yWPrGemNP0zY5DtZdBSxP3XyfrIRXTAzNTDu1s-B9Au7zS91y92n0z2c6Ivr_yFBHhWMzHMVZwLMovUfcMOWhBRkNkmuhNJ2o4sQfs5PmUospdKluTGHpfx4SRwJd-yfgxxUYR0PI_nwteioihbfFcSHNqa6Qq1PJEgZ20pQuZ2XVPRzts0jDLKTFynCQVnw_0lk_Vifx6Kr3Uz_tQ95QJ7_4DOQbIRN2ViWuo_lYSO-5wdWJGVWfm1oQagbt-fV1FunuBhInr88DWBQ63bpBFkpVafZ5GxboiZMYqu_w9tbrkEm2VsLE5V-2uwEXtG8tV75wzIm0UsQsc1pvB1aQP2AI4MUpF0u47GIAWeCvuZdXcQ4Wlbhw-eFrvbhA_x-qjyN3rfsbQWA9ipMVUZpYhXvuIcYaDL3OG1d31JC4eJNI9biOdzXm_fHPkIk6Tl4Zt-r_qOJI8CrbF6KQktvogN7zmyLcF2iN0vWe9ijDLDvmBUjJVZs7yAvKAa-NBYwmmRtFTlpCbf2tLjJyUqRPnzs6sPkWe8BBFadpHbu7TeQBRygpjbXeMeYNB5yjholWclWF26157w9N01GCLomapxRu5FvvSFWjYkrSphcuUpE4Cm.cE5gttNDWTZ2O1vmmP6QzQ/__results___files/__results___6_2.png)\n\n\nHow exactly can we learn segmentation from differently oriented source and mask?",
    "1228057": "Yes, of course it may be different. There is no one correct way to layout most molecules, and that is part of the challenge inherent to this competition. So you can't use the generated image as a mask directly, but perhaps there's other ways it can be of use ;)",
    "1228649": "Perhaps it's possible to segment the synthetic images based on the pre-drawing information (i.e. you will know ahead of time which what will be contained in each part of the image). then once you train a model that learns to segment those synthetic images, you can apply it directly on the images in the training/test set without training any further (i.e. zero-shot transfer learning).",
    "1238103": "I really like the idea of using only the Transformer model, however the ViT paper shows that, although the ViT model produces SOTA results on ImageNet, it requires much more data in order to beat traditional CNN models (eg. ResNet). Since the dataset is only approx. 4M images, unless you generate a very significant amount of additional data, it might be the case that traditional CNN models will work better.",
    "1244412": "Cassava competition had 27000 training images and [the 3rd place solution](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/221150) was an ensemble of ONLY ViT models! I think it's totally doable here with 4M images. \nThe only limitation with ViT is the image resolution. If size matters in this competition then ViT won't be performant since the available pretrained models can't handle more than 384x384 image size.",
    "1260795": "My model has this exact implementation (4.49 best LB score). More precisely I have a slightly modified vit_deit_tiny_distilled_patch16_224 from timm as encoder connected to 3 decoder layers (~30MB model size).\n\nI'm aiming now to increase my model size and improve my preprocessing and data augmentation pipeline.\n\nThe bad thing is that my last model trained during ~75 epochs with 1h per epoch (RTX3090). For inference (autoregressive decoding), it needed around ~22h in float32 precision and ~13h in float16. So basically I can just make one submission a week."
  },
  "source": "meta"
}