{
  "id": 226292,
  "title": "how to simplify the problem",
  "url": "/competitions/bms-molecular-translation/discussion/226292",
  "author_name": "",
  "post_date": "2021-03-16T00:47:55.694339600Z",
  "votes": 17,
  "comment_count": 6,
  "views": 0,
  "content": "<p>the dataset images are created by some python? toolkit and then corrupt with noise.<br>\ni think if we have original images, maybe results would be better.</p>\n<p>one can do the following preprocessing</p>\n<ul>\n<li>detect the orientation and correct it (this is very obvious. the correct orientation should be one that the alphabets are upright)</li>\n</ul>\n<p>there is a trick: width/height ratio</p>\n<ul>\n<li>denoise it. make the letters, digits readable. (or one train skip the denoise step and train with noisty augmented images.)</li>\n<li>normalise to the same scale. i.e. all letters and structure are the same size (or you may want to build a multiscale pyramid model). Note that the scale variation is small, but nevertheless visible</li>\n</ul>\n<p>for those working with OCR software before would know that if the input is of correct size and the image is clear, the results are pretty accurate</p>\n<p>(google for processing methods related to OCR)</p>",
  "messages": [
    {
      "id": "1239727",
      "postDate": "03/16/2021 00:47:55",
      "content": "<p>the dataset images are created by some python? toolkit and then corrupt with noise.<br>\ni think if we have original images, maybe results would be better.</p>\n<p>one can do the following preprocessing</p>\n<ul>\n<li>detect the orientation and correct it (this is very obvious. the correct orientation should be one that the alphabets are upright)</li>\n</ul>\n<p>there is a trick: width/height ratio</p>\n<ul>\n<li>denoise it. make the letters, digits readable. (or one train skip the denoise step and train with noisty augmented images.)</li>\n<li>normalise to the same scale. i.e. all letters and structure are the same size (or you may want to build a multiscale pyramid model). Note that the scale variation is small, but nevertheless visible</li>\n</ul>\n<p>for those working with OCR software before would know that if the input is of correct size and the image is clear, the results are pretty accurate</p>\n<p>(google for processing methods related to OCR)</p>",
      "rawMarkdown": "the dataset images are created by some python? toolkit and then corrupt with noise.\ni think if we have original images, maybe results would be better.\n\none can do the following preprocessing\n- detect the orientation and correct it (this is very obvious. the correct orientation should be one that the alphabets are upright)\n\nthere is a trick: width/height ratio\n\n- denoise it. make the letters, digits readable. (or one train skip the denoise step and train with noisty augmented images.)\n- normalise to the same scale. i.e. all letters and structure are the same size (or you may want to build a multiscale pyramid model). Note that the scale variation is small, but nevertheless visible\n\nfor those working with OCR software before would know that if the input is of correct size and the image is clear, the results are pretty accurate\n\n(google for processing methods related to OCR)",
      "votes": null
    },
    {
      "id": "1239792",
      "postDate": "03/16/2021 02:52:14",
      "content": "<blockquote>\n  <p>for those working with OCR software before would know that if the input is of correct size and the image is clear, the results are pretty accurate</p>\n</blockquote>\n<p>I didn't expect that, the images have a graph representation of InChI notations, how this kind of software can achive accurate results?</p>",
      "rawMarkdown": "> for those working with OCR software before would know that if the input is of correct size and the image is clear, the results are pretty accurate\n\nI didn't expect that, the images have a graph representation of InChI notations, how this kind of software can achive accurate results?",
      "votes": null
    },
    {
      "id": "1239814",
      "postDate": "03/16/2021 03:25:33",
      "content": "<p>no.</p>\n<p>i mean that if you have done development for OCR software, you would know that scale and noise are factors that affect performance.</p>\n<p>the same will apply here. </p>\n<p>it doesn't mean the orc software will work here.</p>\n<p>it means that the factors that affect orc would apply here, i.e. the same factors will also affect performance for molecular structural translation.</p>",
      "rawMarkdown": "no.\n\ni mean that if you have done development for OCR software, you would know that scale and noise are factors that affect performance.\n\nthe same will apply here. \n\nit doesn't mean the orc software will work here.\n\nit means that the factors that affect orc would apply here, i.e. the same factors will also affect performance for molecular structural translation.",
      "votes": null
    },
    {
      "id": "1240320",
      "postDate": "03/16/2021 11:18:42",
      "content": "<p>Thanks for posting your idea. CLAHE could be used to denoise these images. </p>",
      "rawMarkdown": "Thanks for posting your idea. CLAHE could be used to denoise these images.",
      "votes": null
    },
    {
      "id": "1240690",
      "postDate": "03/16/2021 15:22:47",
      "content": "<p>try reverse engineering working backward!</p>\n<ol>\n<li><p>use rdkit to generate chemical structure into multi-channel image<br>\ne.g. one channel for bond, one channel for C, one channel for H, …. , one channel for  atom, the creation of layers should follow your understanding of inchi notation (which itself is in different \"layers\")</p></li>\n<li><p>input these generated channel into a LSTM net (or transformer net)</p></li>\n<li><p>check if you can get high accuracy</p></li>\n</ol>\n<p>This represents the best performance limit. Then think of a way to use computer vision to translate from input images to these layers (e.g. segmentation, faster rcnn, attention, etc)</p>",
      "rawMarkdown": "try reverse engineering working backward!\n\n1. use rdkit to generate chemical structure into multi-channel image\ne.g. one channel for bond, one channel for C, one channel for H, .... , one channel for <start> atom, the creation of layers should follow your understanding of inchi notation (which itself is in different \"layers\")\n\n2. input these generated channel into a LSTM net (or transformer net)\n\n3. check if you can get high accuracy\n\nThis represents the best performance limit. Then think of a way to use computer vision to translate from input images to these layers (e.g. segmentation, faster rcnn, attention, etc)",
      "votes": null
    },
    {
      "id": "1241313",
      "postDate": "03/17/2021 01:57:15",
      "content": "<p>a repeat of what is mentioned here:<br>\n<a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223706\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/223706</a></p>\n<p>why don't you do this:</p>\n<ul>\n<li><p>train a model</p></li>\n<li><p>use the model to predict the test, you end with prediction like<br>\ninchi1,inchi2,inchi3 …inchi1000</p></li>\n<li><p>use rdkit to generate new train image samples for inchi1,inchi2,inchi3 …inchi1000</p></li>\n<li><p>retrain and repeat</p></li>\n</ul>\n<p>note that we have a perfect generator to go from text to image using rdkit. we are doing reverse engineering.<br>\nusing the generator should be part of your solution.</p>\n<p>it is like computer graphics. we have a render to make the image. now we are going from image backward to 3d model.</p>\n<p>data is the key to this competition. we have infinite data. (we have all test and train data). just think of a way to generate and memorize them</p>",
      "rawMarkdown": "a repeat of what is mentioned here:\nhttps://www.kaggle.com/c/bms-molecular-translation/discussion/223706\n\nwhy don't you do this:\n\n- train a model\n\n- use the model to predict the test, you end with prediction like\ninchi1,inchi2,inchi3 …inchi1000\n\n- use rdkit to generate new train image samples for inchi1,inchi2,inchi3 …inchi1000\n\n- retrain and repeat\n\nnote that we have a perfect generator to go from text to image using rdkit. we are doing reverse engineering.\nusing the generator should be part of your solution.\n\nit is like computer graphics. we have a render to make the image. now we are going from image backward to 3d model.\n\ndata is the key to this competition. we have infinite data. (we have all test and train data). just think of a way to generate and memorize them",
      "votes": null
    },
    {
      "id": "1241351",
      "postDate": "03/17/2021 02:24:55",
      "content": "<p>I think that correcting the orientation is non-trivial in this case seeing as many of the letters are symmetric…e.g. <code>H</code>, <code>I</code> and <code>O</code>. Perhaps the easiest way to do this is to re-generate all the images via <code>rdkit</code> and then try to mimic the noise level present in the given training images via augmentations and to then use these images as your primary training images. </p>\n<p>Also, <code>rdkit</code> can produce color-annotated images for the InChI sequences so perhaps we can use some sort of teacher-student training similar to what you proposed in the RANZCR CLiP competition? </p>",
      "rawMarkdown": "I think that correcting the orientation is non-trivial in this case seeing as many of the letters are symmetric...e.g. `H`, `I` and `O`. Perhaps the easiest way to do this is to re-generate all the images via `rdkit` and then try to mimic the noise level present in the given training images via augmentations and to then use these images as your primary training images. \n\nAlso, `rdkit` can produce color-annotated images for the InChI sequences so perhaps we can use some sort of teacher-student training similar to what you proposed in the RANZCR CLiP competition?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1239792,
      "author_name": "hiramcho",
      "author_url": "",
      "post_date": "03/16/2021 02:52:14",
      "content": "<blockquote>\n  <p>for those working with OCR software before would know that if the input is of correct size and the image is clear, the results are pretty accurate</p>\n</blockquote>\n<p>I didn't expect that, the images have a graph representation of InChI notations, how this kind of software can achive accurate results?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1239814,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/16/2021 03:25:33",
          "content": "<p>no.</p>\n<p>i mean that if you have done development for OCR software, you would know that scale and noise are factors that affect performance.</p>\n<p>the same will apply here. </p>\n<p>it doesn't mean the orc software will work here.</p>\n<p>it means that the factors that affect orc would apply here, i.e. the same factors will also affect performance for molecular structural translation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1240320,
      "author_name": "tamilselvanmoorthy",
      "author_url": "",
      "post_date": "03/16/2021 11:18:42",
      "content": "<p>Thanks for posting your idea. CLAHE could be used to denoise these images. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1240690,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/16/2021 15:22:47",
      "content": "<p>try reverse engineering working backward!</p>\n<ol>\n<li><p>use rdkit to generate chemical structure into multi-channel image<br>\ne.g. one channel for bond, one channel for C, one channel for H, …. , one channel for  atom, the creation of layers should follow your understanding of inchi notation (which itself is in different \"layers\")</p></li>\n<li><p>input these generated channel into a LSTM net (or transformer net)</p></li>\n<li><p>check if you can get high accuracy</p></li>\n</ol>\n<p>This represents the best performance limit. Then think of a way to use computer vision to translate from input images to these layers (e.g. segmentation, faster rcnn, attention, etc)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1241313,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/17/2021 01:57:15",
      "content": "<p>a repeat of what is mentioned here:<br>\n<a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223706\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/223706</a></p>\n<p>why don't you do this:</p>\n<ul>\n<li><p>train a model</p></li>\n<li><p>use the model to predict the test, you end with prediction like<br>\ninchi1,inchi2,inchi3 …inchi1000</p></li>\n<li><p>use rdkit to generate new train image samples for inchi1,inchi2,inchi3 …inchi1000</p></li>\n<li><p>retrain and repeat</p></li>\n</ul>\n<p>note that we have a perfect generator to go from text to image using rdkit. we are doing reverse engineering.<br>\nusing the generator should be part of your solution.</p>\n<p>it is like computer graphics. we have a render to make the image. now we are going from image backward to 3d model.</p>\n<p>data is the key to this competition. we have infinite data. (we have all test and train data). just think of a way to generate and memorize them</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1241351,
      "author_name": "tuckerarrants",
      "author_url": "",
      "post_date": "03/17/2021 02:24:55",
      "content": "<p>I think that correcting the orientation is non-trivial in this case seeing as many of the letters are symmetric…e.g. <code>H</code>, <code>I</code> and <code>O</code>. Perhaps the easiest way to do this is to re-generate all the images via <code>rdkit</code> and then try to mimic the noise level present in the given training images via augmentations and to then use these images as your primary training images. </p>\n<p>Also, <code>rdkit</code> can produce color-annotated images for the InChI sequences so perhaps we can use some sort of teacher-student training similar to what you proposed in the RANZCR CLiP competition? </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1239727": "the dataset images are created by some python? toolkit and then corrupt with noise.\ni think if we have original images, maybe results would be better.\n\none can do the following preprocessing\n- detect the orientation and correct it (this is very obvious. the correct orientation should be one that the alphabets are upright)\n\nthere is a trick: width/height ratio\n\n- denoise it. make the letters, digits readable. (or one train skip the denoise step and train with noisty augmented images.)\n- normalise to the same scale. i.e. all letters and structure are the same size (or you may want to build a multiscale pyramid model). Note that the scale variation is small, but nevertheless visible\n\nfor those working with OCR software before would know that if the input is of correct size and the image is clear, the results are pretty accurate\n\n(google for processing methods related to OCR)",
    "1239792": "> for those working with OCR software before would know that if the input is of correct size and the image is clear, the results are pretty accurate\n\nI didn't expect that, the images have a graph representation of InChI notations, how this kind of software can achive accurate results?",
    "1239814": "no.\n\ni mean that if you have done development for OCR software, you would know that scale and noise are factors that affect performance.\n\nthe same will apply here. \n\nit doesn't mean the orc software will work here.\n\nit means that the factors that affect orc would apply here, i.e. the same factors will also affect performance for molecular structural translation.",
    "1240320": "Thanks for posting your idea. CLAHE could be used to denoise these images.",
    "1240690": "try reverse engineering working backward!\n\n1. use rdkit to generate chemical structure into multi-channel image\ne.g. one channel for bond, one channel for C, one channel for H, .... , one channel for <start> atom, the creation of layers should follow your understanding of inchi notation (which itself is in different \"layers\")\n\n2. input these generated channel into a LSTM net (or transformer net)\n\n3. check if you can get high accuracy\n\nThis represents the best performance limit. Then think of a way to use computer vision to translate from input images to these layers (e.g. segmentation, faster rcnn, attention, etc)",
    "1241313": "a repeat of what is mentioned here:\nhttps://www.kaggle.com/c/bms-molecular-translation/discussion/223706\n\nwhy don't you do this:\n\n- train a model\n\n- use the model to predict the test, you end with prediction like\ninchi1,inchi2,inchi3 …inchi1000\n\n- use rdkit to generate new train image samples for inchi1,inchi2,inchi3 …inchi1000\n\n- retrain and repeat\n\nnote that we have a perfect generator to go from text to image using rdkit. we are doing reverse engineering.\nusing the generator should be part of your solution.\n\nit is like computer graphics. we have a render to make the image. now we are going from image backward to 3d model.\n\ndata is the key to this competition. we have infinite data. (we have all test and train data). just think of a way to generate and memorize them",
    "1241351": "I think that correcting the orientation is non-trivial in this case seeing as many of the letters are symmetric...e.g. `H`, `I` and `O`. Perhaps the easiest way to do this is to re-generate all the images via `rdkit` and then try to mimic the noise level present in the given training images via augmentations and to then use these images as your primary training images. \n\nAlso, `rdkit` can produce color-annotated images for the InChI sequences so perhaps we can use some sort of teacher-student training similar to what you proposed in the RANZCR CLiP competition?"
  },
  "source": "meta"
}