{
  "id": 223239,
  "title": "How to solve this problem",
  "url": "/competitions/bms-molecular-translation/discussion/223239",
  "author_name": "ammarali32",
  "post_date": "2021-03-03T03:45:07.222000",
  "votes": 47,
  "comment_count": 17,
  "views": 0,
  "content": "<p>The first step to the solution is to make efficient feature Extraction. The problem that ImageNet weights will not extract appropriate features. That is why I will publish later weights that should perform better. I will attach also the algorithm used to train the feature extractor (i will not be a part of this competition but I hope I will help as much as I could).<br>\nThe second step I am thinking of is using GANS and try to solve this problem as an image captioning problem. <br>\nIf you want also to go in this direction, then these links could be helpful:<br>\n<a href=\"https://github.com/aimagelab/meshed-memory-transformer\" target=\"_blank\">https://github.com/aimagelab/meshed-memory-transformer</a><br>\n<a href=\"https://github.com/LuoweiZhou/VLP\" target=\"_blank\">https://github.com/LuoweiZhou/VLP</a><br>\n<a href=\"https://github.com/rakshithShetty/captionGAN\" target=\"_blank\">https://github.com/rakshithShetty/captionGAN</a><br>\n<a href=\"https://paperswithcode.com/task/image-captioning\" target=\"_blank\">https://paperswithcode.com/task/image-captioning</a><br>\nThe starter I already published is using Encoder and decoder for the features will another loss function small changes could also give good results but in general, it is not a good solution.</p>\n<p>IMPORTANT NOTES:<br>\nDefining rules for the labels will help a lot and I can see directly a lot of rules that will make the solution less complicated.<br>\nI see this as a Transformers, GANS competition.<br>\nHope that will be useful. Best of luck to all of you</p>\n<p>UPDATE:</p>\n<ul>\n<li>Trained model to get the molecular smiles Representation<br>\n<a href=\"https://github.com/Kohulan/DECIMER-Image-to-SMILES\" target=\"_blank\">https://github.com/Kohulan/DECIMER-Image-to-SMILES</a> </li>\n<li>Open source library to convert from SMILES to InChi<br>\n<a href=\"https://pypi.org/project/ChemSpiPy/\" target=\"_blank\">https://pypi.org/project/ChemSpiPy/</a></li>\n</ul>",
  "messages": [
    {
      "id": 1224767,
      "postDate": "2021-03-03T03:45:07.223Z",
      "content": "<p>The first step to the solution is to make efficient feature Extraction. The problem that ImageNet weights will not extract appropriate features. That is why I will publish later weights that should perform better. I will attach also the algorithm used to train the feature extractor (i will not be a part of this competition but I hope I will help as much as I could).<br>\nThe second step I am thinking of is using GANS and try to solve this problem as an image captioning problem. <br>\nIf you want also to go in this direction, then these links could be helpful:<br>\n<a href=\"https://github.com/aimagelab/meshed-memory-transformer\" target=\"_blank\">https://github.com/aimagelab/meshed-memory-transformer</a><br>\n<a href=\"https://github.com/LuoweiZhou/VLP\" target=\"_blank\">https://github.com/LuoweiZhou/VLP</a><br>\n<a href=\"https://github.com/rakshithShetty/captionGAN\" target=\"_blank\">https://github.com/rakshithShetty/captionGAN</a><br>\n<a href=\"https://paperswithcode.com/task/image-captioning\" target=\"_blank\">https://paperswithcode.com/task/image-captioning</a><br>\nThe starter I already published is using Encoder and decoder for the features will another loss function small changes could also give good results but in general, it is not a good solution.</p>\n<p>IMPORTANT NOTES:<br>\nDefining rules for the labels will help a lot and I can see directly a lot of rules that will make the solution less complicated.<br>\nI see this as a Transformers, GANS competition.<br>\nHope that will be useful. Best of luck to all of you</p>\n<p>UPDATE:</p>\n<ul>\n<li>Trained model to get the molecular smiles Representation<br>\n<a href=\"https://github.com/Kohulan/DECIMER-Image-to-SMILES\" target=\"_blank\">https://github.com/Kohulan/DECIMER-Image-to-SMILES</a> </li>\n<li>Open source library to convert from SMILES to InChi<br>\n<a href=\"https://pypi.org/project/ChemSpiPy/\" target=\"_blank\">https://pypi.org/project/ChemSpiPy/</a></li>\n</ul>",
      "rawMarkdown": "The first step to the solution is to make efficient feature Extraction. The problem that ImageNet weights will not extract appropriate features. That is why I will publish later weights that should perform better. I will attach also the algorithm used to train the feature extractor (i will not be a part of this competition but I hope I will help as much as I could).\nThe second step I am thinking of is using GANS and try to solve this problem as an image captioning problem. \nIf you want also to go in this direction, then these links could be helpful:\nhttps://github.com/aimagelab/meshed-memory-transformer\nhttps://github.com/LuoweiZhou/VLP\nhttps://github.com/rakshithShetty/captionGAN\nhttps://paperswithcode.com/task/image-captioning\nThe starter I already published is using Encoder and decoder for the features will another loss function small changes could also give good results but in general, it is not a good solution.\n\nIMPORTANT NOTES:\nDefining rules for the labels will help a lot and I can see directly a lot of rules that will make the solution less complicated.\nI see this as a Transformers, GANS competition.\nHope that will be useful. Best of luck to all of you\n\nUPDATE:\n- Trained model to get the molecular smiles Representation\nhttps://github.com/Kohulan/DECIMER-Image-to-SMILES \n- Open source library to convert from SMILES to InChi\nhttps://pypi.org/project/ChemSpiPy/\n",
      "votes": 46
    },
    {
      "id": 1244699,
      "postDate": "2021-03-19T06:55:55.687Z",
      "content": "<p>I understand the role of transformer in this competition. What is the role of GAN? </p>",
      "rawMarkdown": "I understand the role of transformer in this competition. What is the role of GAN? ",
      "votes": 1,
      "replies": [
        {
          "id": 1244710,
          "postDate": "2021-03-19T07:11:45.580Z",
          "content": "<p>I will refer to these papers:<br>\n<a href=\"https://ieeexplore.ieee.org/document/8545049\" target=\"_blank\">https://ieeexplore.ieee.org/document/8545049</a><br>\n<a href=\"https://www.researchgate.net/publication/335699458_Conditional_GANs_for_Image_Captioning_with_Sentiments\" target=\"_blank\">https://www.researchgate.net/publication/335699458_Conditional_GANs_for_Image_Captioning_with_Sentiments</a><br>\n<a href=\"https://vigilworkshop.github.io/static/papers/38.pdf\" target=\"_blank\">https://vigilworkshop.github.io/static/papers/38.pdf</a></p>",
          "rawMarkdown": "I will refer to these papers:\nhttps://ieeexplore.ieee.org/document/8545049\nhttps://www.researchgate.net/publication/335699458_Conditional_GANs_for_Image_Captioning_with_Sentiments\nhttps://vigilworkshop.github.io/static/papers/38.pdf",
          "votes": 1
        },
        {
          "id": 1244724,
          "postDate": "2021-03-19T07:23:50.053Z",
          "content": "<p>Wow! Interesting! This competition just got exciting for me.</p>\n<p>Never really thought generative models can be a major ingredient in a ML competition.</p>",
          "rawMarkdown": "Wow! Interesting! This competition just got exciting for me.\n\nNever really thought generative models can be a major ingredient in a ML competition.",
          "votes": 2
        },
        {
          "id": 1244728,
          "postDate": "2021-03-19T07:30:50.033Z",
          "content": "<p>Yes, it is a really interesting competition. Unfortunately, 4 Million images need high computational resources. Good luck and keep the great work  </p>",
          "rawMarkdown": "Yes, it is a really interesting competition. Unfortunately, 4 Million images need high computational resources. Good luck and keep the great work  "
        },
        {
          "id": 1245194,
          "postDate": "2021-03-19T14:55:59.780Z",
          "content": "<p>Yea, training a generative model on 4-million images will need quite a bit of compute time.</p>\n<p>On the flip side, unlike RANZCR where high res and large models are keys, I don't think you need a particularly large model (large model wouldn't hurt though) for this. </p>\n<p>I may revisit this after my thesis submission and presentation.</p>",
          "rawMarkdown": "Yea, training a generative model on 4-million images will need quite a bit of compute time.\n\nOn the flip side, unlike RANZCR where high res and large models are keys, I don't think you need a particularly large model (large model wouldn't hurt though) for this. \n\nI may revisit this after my thesis submission and presentation."
        }
      ]
    },
    {
      "id": 1235154,
      "postDate": "2021-03-11T21:54:37.550Z",
      "content": "<p><a href=\"https://pypi.org/project/ChemSpiPy/\" target=\"_blank\">https://pypi.org/project/ChemSpiPy/</a> is limited to 1000 calls per month…</p>",
      "rawMarkdown": "https://pypi.org/project/ChemSpiPy/ is limited to 1000 calls per month...",
      "votes": 1,
      "replies": [
        {
          "id": 1244708,
          "postDate": "2021-03-19T07:07:15.347Z",
          "content": "<p>I am sorry i didn't know that </p>",
          "rawMarkdown": "I am sorry i didn't know that "
        }
      ]
    },
    {
      "id": 1231063,
      "postDate": "2021-03-08T16:34:17.993Z",
      "content": "<p><a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a> Why will you not be part of this competition ?</p>",
      "rawMarkdown": "@ammarali32 Why will you not be part of this competition ?",
      "votes": 1,
      "replies": [
        {
          "id": 1231586,
          "postDate": "2021-03-09T05:00:54.153Z",
          "content": "<p>Thanks for your interest. Actually, I don't have time or resources for large data I am just using Kaggle resources ))</p>",
          "rawMarkdown": "Thanks for your interest. Actually, I don't have time or resources for large data I am just using Kaggle resources ))",
          "votes": 1
        },
        {
          "id": 1246302,
          "postDate": "2021-03-20T16:19:25.487Z",
          "content": "<p>Are you still using only Kaggle resources? If yes, then I am really curious to know how did you achieve this score because I don't have resources either other than Kaggle. <br>\n<a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a> </p>",
          "rawMarkdown": "Are you still using only Kaggle resources? If yes, then I am really curious to know how did you achieve this score because I don't have resources either other than Kaggle. \n@ammarali32 ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1230684,
      "postDate": "2021-03-08T11:04:45.247Z",
      "content": "<p>Thanks for posting! </p>\n<blockquote>\n  <p>Defining rules for the labels will help a lot and I can see directly a lot of rules that will make the solution less complicated.</p>\n</blockquote>\n<p>Sorry it wasn't clear to me, are you referring to writing a fn that would correct the o/p from models to correct the notation  generated based on what molecules are actually possible?</p>",
      "rawMarkdown": "Thanks for posting! \n\n> Defining rules for the labels will help a lot and I can see directly a lot of rules that will make the solution less complicated.\n\nSorry it wasn't clear to me, are you referring to writing a fn that would correct the o/p from models to correct the notation  generated based on what molecules are actually possible?",
      "votes": 1,
      "replies": [
        {
          "id": 1231584,
          "postDate": "2021-03-09T04:59:25.617Z",
          "content": "<p>Not exactly but for example assume that you are able with high accuracy to predict the number of carbon atoms and the other atoms is already written in this case you can check the number of Hydrogen atoms in \"chemistry\" mathematically. This presentation \"InChi\" describes the distribution and relations for a particular molecular. you need to know how to write InChi representation manually after that a lot of rules for auto-correction will help.</p>",
          "rawMarkdown": "Not exactly but for example assume that you are able with high accuracy to predict the number of carbon atoms and the other atoms is already written in this case you can check the number of Hydrogen atoms in \"chemistry\" mathematically. This presentation \"InChi\" describes the distribution and relations for a particular molecular. you need to know how to write InChi representation manually after that a lot of rules for auto-correction will help.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1225048,
      "postDate": "2021-03-03T09:14:37.803Z",
      "content": "<p>Interesting review from a Data Polymath</p>",
      "rawMarkdown": "Interesting review from a Data Polymath",
      "votes": 2,
      "replies": [
        {
          "id": 1225054,
          "postDate": "2021-03-03T09:16:07.050Z",
          "content": "<p>Thank you))</p>",
          "rawMarkdown": "Thank you))",
          "votes": 2
        },
        {
          "id": 1231062,
          "postDate": "2021-03-08T16:33:26.987Z",
          "content": "<p><a href=\"https://www.kaggle.com/kingabzpro\" target=\"_blank\">@kingabzpro</a> Looks like you changed the color of your profile pic from red to black .</p>",
          "rawMarkdown": "@kingabzpro Looks like you changed the color of your profile pic from red to black .",
          "votes": 1
        },
        {
          "id": 1231070,
          "postDate": "2021-03-08T16:39:02.313Z",
          "content": "<p>yes, just like Kaggle changed from color to black, just trying to blend in. </p>",
          "rawMarkdown": "yes, just like Kaggle changed from color to black, just trying to blend in. "
        }
      ]
    },
    {
      "id": 1227077,
      "postDate": "2021-03-05T07:09:04.237Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1244699,
      "author_name": "JunYong Tong",
      "author_url": "",
      "post_date": "2021-03-19T06:55:55.687000",
      "content": "<p>I understand the role of transformer in this competition. What is the role of GAN? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1244710,
          "author_name": "ammarali32",
          "author_url": "",
          "post_date": "2021-03-19T07:11:45.580000",
          "content": "<p>I will refer to these papers:<br>\n<a href=\"https://ieeexplore.ieee.org/document/8545049\" target=\"_blank\">https://ieeexplore.ieee.org/document/8545049</a><br>\n<a href=\"https://www.researchgate.net/publication/335699458_Conditional_GANs_for_Image_Captioning_with_Sentiments\" target=\"_blank\">https://www.researchgate.net/publication/335699458_Conditional_GANs_for_Image_Captioning_with_Sentiments</a><br>\n<a href=\"https://vigilworkshop.github.io/static/papers/38.pdf\" target=\"_blank\">https://vigilworkshop.github.io/static/papers/38.pdf</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1244724,
          "author_name": "JunYong Tong",
          "author_url": "",
          "post_date": "2021-03-19T07:23:50.053000",
          "content": "<p>Wow! Interesting! This competition just got exciting for me.</p>\n<p>Never really thought generative models can be a major ingredient in a ML competition.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1244728,
          "author_name": "ammarali32",
          "author_url": "",
          "post_date": "2021-03-19T07:30:50.033000",
          "content": "<p>Yes, it is a really interesting competition. Unfortunately, 4 Million images need high computational resources. Good luck and keep the great work  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1245194,
          "author_name": "JunYong Tong",
          "author_url": "",
          "post_date": "2021-03-19T14:55:59.780000",
          "content": "<p>Yea, training a generative model on 4-million images will need quite a bit of compute time.</p>\n<p>On the flip side, unlike RANZCR where high res and large models are keys, I don't think you need a particularly large model (large model wouldn't hurt though) for this. </p>\n<p>I may revisit this after my thesis submission and presentation.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1235154,
      "author_name": "Alexander Scarlat MD",
      "author_url": "",
      "post_date": "2021-03-11T21:54:37.550000",
      "content": "<p><a href=\"https://pypi.org/project/ChemSpiPy/\" target=\"_blank\">https://pypi.org/project/ChemSpiPy/</a> is limited to 1000 calls per month…</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1244708,
          "author_name": "ammarali32",
          "author_url": "",
          "post_date": "2021-03-19T07:07:15.347000",
          "content": "<p>I am sorry i didn't know that </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1231063,
      "author_name": "Tensor Girl",
      "author_url": "",
      "post_date": "2021-03-08T16:34:17.993000",
      "content": "<p><a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a> Why will you not be part of this competition ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1231586,
          "author_name": "ammarali32",
          "author_url": "",
          "post_date": "2021-03-09T05:00:54.153000",
          "content": "<p>Thanks for your interest. Actually, I don't have time or resources for large data I am just using Kaggle resources ))</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1246302,
          "author_name": "Shweta Goyal",
          "author_url": "",
          "post_date": "2021-03-20T16:19:25.487000",
          "content": "<p>Are you still using only Kaggle resources? If yes, then I am really curious to know how did you achieve this score because I don't have resources either other than Kaggle. <br>\n<a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1230684,
      "author_name": "Sanyam Bhutani",
      "author_url": "",
      "post_date": "2021-03-08T11:04:45.247000",
      "content": "<p>Thanks for posting! </p>\n<blockquote>\n  <p>Defining rules for the labels will help a lot and I can see directly a lot of rules that will make the solution less complicated.</p>\n</blockquote>\n<p>Sorry it wasn't clear to me, are you referring to writing a fn that would correct the o/p from models to correct the notation  generated based on what molecules are actually possible?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1231584,
          "author_name": "ammarali32",
          "author_url": "",
          "post_date": "2021-03-09T04:59:25.617000",
          "content": "<p>Not exactly but for example assume that you are able with high accuracy to predict the number of carbon atoms and the other atoms is already written in this case you can check the number of Hydrogen atoms in \"chemistry\" mathematically. This presentation \"InChi\" describes the distribution and relations for a particular molecular. you need to know how to write InChi representation manually after that a lot of rules for auto-correction will help.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1225048,
      "author_name": "Abid Ali Awan",
      "author_url": "",
      "post_date": "2021-03-03T09:14:37.803000",
      "content": "<p>Interesting review from a Data Polymath</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1225054,
          "author_name": "ammarali32",
          "author_url": "",
          "post_date": "2021-03-03T09:16:07.050000",
          "content": "<p>Thank you))</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1231062,
          "author_name": "Tensor Girl",
          "author_url": "",
          "post_date": "2021-03-08T16:33:26.987000",
          "content": "<p><a href=\"https://www.kaggle.com/kingabzpro\" target=\"_blank\">@kingabzpro</a> Looks like you changed the color of your profile pic from red to black .</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1231070,
          "author_name": "Abid Ali Awan",
          "author_url": "",
          "post_date": "2021-03-08T16:39:02.313000",
          "content": "<p>yes, just like Kaggle changed from color to black, just trying to blend in. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1227077,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-05T07:09:04.237000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1224767": "The first step to the solution is to make efficient feature Extraction. The problem that ImageNet weights will not extract appropriate features. That is why I will publish later weights that should perform better. I will attach also the algorithm used to train the feature extractor (i will not be a part of this competition but I hope I will help as much as I could).\nThe second step I am thinking of is using GANS and try to solve this problem as an image captioning problem. \nIf you want also to go in this direction, then these links could be helpful:\nhttps://github.com/aimagelab/meshed-memory-transformer\nhttps://github.com/LuoweiZhou/VLP\nhttps://github.com/rakshithShetty/captionGAN\nhttps://paperswithcode.com/task/image-captioning\nThe starter I already published is using Encoder and decoder for the features will another loss function small changes could also give good results but in general, it is not a good solution.\n\nIMPORTANT NOTES:\nDefining rules for the labels will help a lot and I can see directly a lot of rules that will make the solution less complicated.\nI see this as a Transformers, GANS competition.\nHope that will be useful. Best of luck to all of you\n\nUPDATE:\n- Trained model to get the molecular smiles Representation\nhttps://github.com/Kohulan/DECIMER-Image-to-SMILES \n- Open source library to convert from SMILES to InChi\nhttps://pypi.org/project/ChemSpiPy/\n",
    "1244699": "I understand the role of transformer in this competition. What is the role of GAN? ",
    "1235154": "https://pypi.org/project/ChemSpiPy/ is limited to 1000 calls per month...",
    "1231063": "@ammarali32 Why will you not be part of this competition ?",
    "1230684": "Thanks for posting! \n\n> Defining rules for the labels will help a lot and I can see directly a lot of rules that will make the solution less complicated.\n\nSorry it wasn't clear to me, are you referring to writing a fn that would correct the o/p from models to correct the notation  generated based on what molecules are actually possible?",
    "1225048": "Interesting review from a Data Polymath",
    "1227077": ""
  }
}