{
  "id": 228139,
  "title": "Chemical Formula Prediction as Multi-Output Regression",
  "url": "/competitions/bms-molecular-translation/discussion/228139",
  "author_name": "",
  "post_date": "2021-03-23T14:36:35.518370400Z",
  "votes": 22,
  "comment_count": 8,
  "views": 0,
  "content": "<p>One of the difficulties of this competition is that the prediction target <code>InChI</code> has <strong>variable length</strong>.</p>\n<p>When we split <code>InChI</code> by <code>/</code>,  the second part is  <strong>chemical formula</strong> which represents the number of atoms in each molecular. As the kind of atoms in training data is limited to 12(<code>B</code>, <code>Br</code>, <code>C</code>, <code>Cl</code>, <code>F</code>, <code>H</code>, <code>I</code>, <code>N</code>, <code>O</code>, <code>P</code>, <code>S</code>, and <code>Si</code>), we can represents a chemical formula by a <strong>fixed length</strong> vector and treat chemical formula prediction task as <strong>multi-output regression task</strong>.</p>\n<p>I've published two notebooks about Chemical Formula Prediction.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/ttahara/bms-mt-chemical-formula-regression-training\" target=\"_blank\">BMS-MT: Chemical Formula Regression [Training]</a></li>\n<li><a href=\"https://www.kaggle.com/ttahara/bms-mt-chemical-formula-regression-inference\" target=\"_blank\">BMS-MT: Chemical Formula Regression [Inference]</a></li>\n</ul>\n<p>OOF Levenshtein distance for chemical formula is <strong>0.9503</strong> although I used <strong>only 4%</strong> of training data. The more data you use, the better score you can achieve.</p>\n<p>In inference notebook, I tried solving the competition task by clustering approach using the predicted number of atoms as feature. This achieved slightly better (LB: 67.70) than other naive baselines.</p>\n<h2>What's next ?</h2>\n<p>Here is local Levenshtein distance score of each part of <code>InChI</code> by regression and clustering approach:</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>parts</th>\n<th>mean_LS</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>InChI_0</td>\n<td>0.000000e+00</td>\n</tr>\n<tr>\n<td>1</td>\n<td>InChI_1</td>\n<td>9.503252e-01</td>\n</tr>\n<tr>\n<td>2</td>\n<td>InChI_2</td>\n<td><strong>4.455307e+01</strong></td>\n</tr>\n<tr>\n<td>3</td>\n<td>InChI_3</td>\n<td><strong>1.762780e+01</strong></td>\n</tr>\n<tr>\n<td>4</td>\n<td>InChI_4</td>\n<td>1.546733e+00</td>\n</tr>\n<tr>\n<td>5</td>\n<td>InChI_5</td>\n<td>3.841467e-01</td>\n</tr>\n<tr>\n<td>6</td>\n<td>InChI_6</td>\n<td>3.256986e-01</td>\n</tr>\n<tr>\n<td>7</td>\n<td>InChI_7</td>\n<td>2.150248e-02</td>\n</tr>\n<tr>\n<td>8</td>\n<td>InChI_8</td>\n<td>4.479854e-04</td>\n</tr>\n<tr>\n<td>9</td>\n<td>InChI_9</td>\n<td>3.011320e-05</td>\n</tr>\n<tr>\n<td>10</td>\n<td>InChI_10</td>\n<td>8.250192e-07</td>\n</tr>\n<tr>\n<td>11</td>\n<td>InChI</td>\n<td>6.472638e+01</td>\n</tr>\n</tbody>\n</table>\n<p>As you see, <code>InChI_2</code> (atom connections information)  and <code>InChI_3</code> (hydrogen atoms information) are the majority of the Levenshtein distance of whole the <code>InChI</code>. To focus on these, I'll try <strong>multi-task approach</strong>: multi-output regression for <code>InChI_1</code> and sequence prediction for <code>InChI_2</code> and <code>InChI_3</code>.</p>\n<p>I am very curious about which is better, multi-task approach or single sequence prediction approach  such as <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> 's <a href=\"https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-inference\" target=\"_blank\">baseline</a>.</p>",
  "messages": [
    {
      "id": "1249812",
      "postDate": "03/23/2021 14:36:35",
      "content": "<p>One of the difficulties of this competition is that the prediction target <code>InChI</code> has <strong>variable length</strong>.</p>\n<p>When we split <code>InChI</code> by <code>/</code>,  the second part is  <strong>chemical formula</strong> which represents the number of atoms in each molecular. As the kind of atoms in training data is limited to 12(<code>B</code>, <code>Br</code>, <code>C</code>, <code>Cl</code>, <code>F</code>, <code>H</code>, <code>I</code>, <code>N</code>, <code>O</code>, <code>P</code>, <code>S</code>, and <code>Si</code>), we can represents a chemical formula by a <strong>fixed length</strong> vector and treat chemical formula prediction task as <strong>multi-output regression task</strong>.</p>\n<p>I've published two notebooks about Chemical Formula Prediction.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/ttahara/bms-mt-chemical-formula-regression-training\" target=\"_blank\">BMS-MT: Chemical Formula Regression [Training]</a></li>\n<li><a href=\"https://www.kaggle.com/ttahara/bms-mt-chemical-formula-regression-inference\" target=\"_blank\">BMS-MT: Chemical Formula Regression [Inference]</a></li>\n</ul>\n<p>OOF Levenshtein distance for chemical formula is <strong>0.9503</strong> although I used <strong>only 4%</strong> of training data. The more data you use, the better score you can achieve.</p>\n<p>In inference notebook, I tried solving the competition task by clustering approach using the predicted number of atoms as feature. This achieved slightly better (LB: 67.70) than other naive baselines.</p>\n<h2>What's next ?</h2>\n<p>Here is local Levenshtein distance score of each part of <code>InChI</code> by regression and clustering approach:</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>parts</th>\n<th>mean_LS</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>InChI_0</td>\n<td>0.000000e+00</td>\n</tr>\n<tr>\n<td>1</td>\n<td>InChI_1</td>\n<td>9.503252e-01</td>\n</tr>\n<tr>\n<td>2</td>\n<td>InChI_2</td>\n<td><strong>4.455307e+01</strong></td>\n</tr>\n<tr>\n<td>3</td>\n<td>InChI_3</td>\n<td><strong>1.762780e+01</strong></td>\n</tr>\n<tr>\n<td>4</td>\n<td>InChI_4</td>\n<td>1.546733e+00</td>\n</tr>\n<tr>\n<td>5</td>\n<td>InChI_5</td>\n<td>3.841467e-01</td>\n</tr>\n<tr>\n<td>6</td>\n<td>InChI_6</td>\n<td>3.256986e-01</td>\n</tr>\n<tr>\n<td>7</td>\n<td>InChI_7</td>\n<td>2.150248e-02</td>\n</tr>\n<tr>\n<td>8</td>\n<td>InChI_8</td>\n<td>4.479854e-04</td>\n</tr>\n<tr>\n<td>9</td>\n<td>InChI_9</td>\n<td>3.011320e-05</td>\n</tr>\n<tr>\n<td>10</td>\n<td>InChI_10</td>\n<td>8.250192e-07</td>\n</tr>\n<tr>\n<td>11</td>\n<td>InChI</td>\n<td>6.472638e+01</td>\n</tr>\n</tbody>\n</table>\n<p>As you see, <code>InChI_2</code> (atom connections information)  and <code>InChI_3</code> (hydrogen atoms information) are the majority of the Levenshtein distance of whole the <code>InChI</code>. To focus on these, I'll try <strong>multi-task approach</strong>: multi-output regression for <code>InChI_1</code> and sequence prediction for <code>InChI_2</code> and <code>InChI_3</code>.</p>\n<p>I am very curious about which is better, multi-task approach or single sequence prediction approach  such as <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> 's <a href=\"https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-inference\" target=\"_blank\">baseline</a>.</p>",
      "rawMarkdown": "One of the difficulties of this competition is that the prediction target `InChI` has **variable length**.\n\nWhen we split `InChI` by `/`,  the second part is  **chemical formula** which represents the number of atoms in each molecular. As the kind of atoms in training data is limited to 12(`B`, `Br`, `C`, `Cl`, `F`, `H`, `I`, `N`, `O`, `P`, `S`, and `Si`), we can represents a chemical formula by a **fixed length** vector and treat chemical formula prediction task as **multi-output regression task**.\n\nI've published two notebooks about Chemical Formula Prediction.\n\n* [BMS-MT: Chemical Formula Regression [Training]](https://www.kaggle.com/ttahara/bms-mt-chemical-formula-regression-training)\n* [BMS-MT: Chemical Formula Regression [Inference]](https://www.kaggle.com/ttahara/bms-mt-chemical-formula-regression-inference)\n\nOOF Levenshtein distance for chemical formula is **0.9503** although I used **only 4%** of training data. The more data you use, the better score you can achieve.\n\nIn inference notebook, I tried solving the competition task by clustering approach using the predicted number of atoms as feature. This achieved slightly better (LB: 67.70) than other naive baselines.\n\n## What's next ?\n\nHere is local Levenshtein distance score of each part of `InChI` by regression and clustering approach:\n\n|     |   parts  |      mean_LS |\n|:---:|:--------:|:------------:|\n|0    | InChI_0  | 0.000000e+00 |\n|1    | InChI_1  | 9.503252e-01 |\n|2    | InChI_2  | **4.455307e+01** |\n|3    | InChI_3  | **1.762780e+01** |\n|4    | InChI_4  | 1.546733e+00 |\n|5    | InChI_5  | 3.841467e-01 |\n|6    | InChI_6  | 3.256986e-01 |\n|7    | InChI_7  | 2.150248e-02 |\n|8    | InChI_8  | 4.479854e-04 |\n|9    | InChI_9  | 3.011320e-05 |\n| 10  | InChI_10 | 8.250192e-07 |\n| 11  |   InChI  | 6.472638e+01 |\n\nAs you see, `InChI_2` (atom connections information)  and `InChI_3` (hydrogen atoms information) are the majority of the Levenshtein distance of whole the `InChI`. To focus on these, I'll try **multi-task approach**: multi-output regression for `InChI_1` and sequence prediction for `InChI_2` and `InChI_3`.\n\nI am very curious about which is better, multi-task approach or single sequence prediction approach  such as @yasufuminakama 's [baseline](https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-inference).",
      "votes": null
    },
    {
      "id": "1251771",
      "postDate": "03/25/2021 05:44:41",
      "content": "<p>Finally a new kind of approach (it was getting boring as I couldn't think of any new); Thanks</p>",
      "rawMarkdown": "Finally a new kind of approach (it was getting boring as I couldn't think of any new); Thanks",
      "votes": null
    },
    {
      "id": "1252002",
      "postDate": "03/25/2021 10:34:16",
      "content": "<p>I think deep understanding of <code>InChI</code> leads to various approaches other than single sequence predicton by encoder-decoder model.</p>\n<p>I hope there wil be a wide variety of solutions at the end of  competition :)</p>",
      "rawMarkdown": "I think deep understanding of `InChI` leads to various approaches other than single sequence predicton by encoder-decoder model.\n\nI hope there wil be a wide variety of solutions at the end of  competition :)",
      "votes": null
    },
    {
      "id": "1255352",
      "postDate": "03/28/2021 17:53:40",
      "content": "<p>The input is an image, which can be represented as a vector. Theoretically, the output is a different representation of the same thing, so it should also be convertible to a vector. The trick is to convert InChI into a vector that can be converted back to InChI.</p>",
      "rawMarkdown": "The input is an image, which can be represented as a vector. Theoretically, the output is a different representation of the same thing, so it should also be convertible to a vector. The trick is to convert InChI into a vector that can be converted back to InChI.",
      "votes": null
    },
    {
      "id": "1257210",
      "postDate": "03/30/2021 16:12:47",
      "content": "<p>Are you saying some sort of Autoencoder can do it?</p>",
      "rawMarkdown": "Are you saying some sort of Autoencoder can do it?",
      "votes": null
    },
    {
      "id": "1257212",
      "postDate": "03/30/2021 16:13:04",
      "content": "<p>Are you saying some sort of Autoencoder can do it?</p>",
      "rawMarkdown": "Are you saying some sort of Autoencoder can do it?",
      "votes": null
    },
    {
      "id": "1258180",
      "postDate": "03/31/2021 12:09:54",
      "content": "<p>two separate ideas here:</p>\n<ul>\n<li><p>think of each /c, /h, /m, etc as a query. <br>\ngiven \"image + /x\", we are to retrive the structure for \"x\"</p></li>\n<li><p>divide the substring of  InchI to words (i.e. entries of a dictionary). You can build a classifier to classifiy if the image contains the word or not. i.e. multi-class classification problem.</p></li>\n<li><p>given a set of \"detected words\", predict the best arrangment</p></li>\n</ul>\n<hr>\n<p>i wonder if anyone can suggest a good tokonizer for chemical InChI words?</p>",
      "rawMarkdown": "two separate ideas here:\n\n- think of each /c, /h, /m, etc as a query. \ngiven \"image + /x\", we are to retrive the structure for \"x\"\n\n- divide the substring of  InchI to words (i.e. entries of a dictionary). You can build a classifier to classifiy if the image contains the word or not. i.e. multi-class classification problem.\n\n- given a set of \"detected words\", predict the best arrangment\n\n---\n\ni wonder if anyone can suggest a good tokonizer for chemical InChI words?",
      "votes": null
    },
    {
      "id": "1258389",
      "postDate": "03/31/2021 15:11:26",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/affjljoo3581/bms-molecular-translation-train-inchi-tokenizer\" target=\"_blank\">WordPiece tokenizer for InChI</a> kernel using AllenNLP</li>\n<li><a href=\"https://pypi.org/project/tokenizers/\" target=\"_blank\">BPE and others</a></li>\n<li><a href=\"https://github.com/huggingface/tokenizers\" target=\"_blank\">HuggingFace tokenizers</a></li>\n</ul>",
      "rawMarkdown": "hengck23\n- [WordPiece tokenizer for InChI](https://www.kaggle.com/affjljoo3581/bms-molecular-translation-train-inchi-tokenizer) kernel using AllenNLP\n- [BPE and others](https://pypi.org/project/tokenizers/)\n- [HuggingFace tokenizers](https://github.com/huggingface/tokenizers)",
      "votes": null
    },
    {
      "id": "1258633",
      "postDate": "03/31/2021 19:04:21",
      "content": "<p>i wonder is there any specialised one for InCHI, e.g.</p>\n<p><img src=\"https://github.com/XinhaoLi74/SmilesPE/raw/master/TOC.PNG\" alt=\"\"><br>\n<a href=\"https://github.com/XinhaoLi74/SmilesPE\" target=\"_blank\">https://github.com/XinhaoLi74/SmilesPE</a></p>",
      "rawMarkdown": "i wonder is there any specialised one for InCHI, e.g.\n\n![](https://github.com/XinhaoLi74/SmilesPE/raw/master/TOC.PNG)\nhttps://github.com/XinhaoLi74/SmilesPE",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1251771,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "03/25/2021 05:44:41",
      "content": "<p>Finally a new kind of approach (it was getting boring as I couldn't think of any new); Thanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 1252002,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "03/25/2021 10:34:16",
          "content": "<p>I think deep understanding of <code>InChI</code> leads to various approaches other than single sequence predicton by encoder-decoder model.</p>\n<p>I hope there wil be a wide variety of solutions at the end of  competition :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1255352,
      "author_name": "solorzano",
      "author_url": "",
      "post_date": "03/28/2021 17:53:40",
      "content": "<p>The input is an image, which can be represented as a vector. Theoretically, the output is a different representation of the same thing, so it should also be convertible to a vector. The trick is to convert InChI into a vector that can be converted back to InChI.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1257210,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "03/30/2021 16:12:47",
          "content": "<p>Are you saying some sort of Autoencoder can do it?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1257212,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "03/30/2021 16:13:04",
          "content": "<p>Are you saying some sort of Autoencoder can do it?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1258180,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/31/2021 12:09:54",
      "content": "<p>two separate ideas here:</p>\n<ul>\n<li><p>think of each /c, /h, /m, etc as a query. <br>\ngiven \"image + /x\", we are to retrive the structure for \"x\"</p></li>\n<li><p>divide the substring of  InchI to words (i.e. entries of a dictionary). You can build a classifier to classifiy if the image contains the word or not. i.e. multi-class classification problem.</p></li>\n<li><p>given a set of \"detected words\", predict the best arrangment</p></li>\n</ul>\n<hr>\n<p>i wonder if anyone can suggest a good tokonizer for chemical InChI words?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1258389,
          "author_name": "sirishks",
          "author_url": "",
          "post_date": "03/31/2021 15:11:26",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/affjljoo3581/bms-molecular-translation-train-inchi-tokenizer\" target=\"_blank\">WordPiece tokenizer for InChI</a> kernel using AllenNLP</li>\n<li><a href=\"https://pypi.org/project/tokenizers/\" target=\"_blank\">BPE and others</a></li>\n<li><a href=\"https://github.com/huggingface/tokenizers\" target=\"_blank\">HuggingFace tokenizers</a></li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1258633,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/31/2021 19:04:21",
          "content": "<p>i wonder is there any specialised one for InCHI, e.g.</p>\n<p><img src=\"https://github.com/XinhaoLi74/SmilesPE/raw/master/TOC.PNG\" alt=\"\"><br>\n<a href=\"https://github.com/XinhaoLi74/SmilesPE\" target=\"_blank\">https://github.com/XinhaoLi74/SmilesPE</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1249812": "One of the difficulties of this competition is that the prediction target `InChI` has **variable length**.\n\nWhen we split `InChI` by `/`,  the second part is  **chemical formula** which represents the number of atoms in each molecular. As the kind of atoms in training data is limited to 12(`B`, `Br`, `C`, `Cl`, `F`, `H`, `I`, `N`, `O`, `P`, `S`, and `Si`), we can represents a chemical formula by a **fixed length** vector and treat chemical formula prediction task as **multi-output regression task**.\n\nI've published two notebooks about Chemical Formula Prediction.\n\n* [BMS-MT: Chemical Formula Regression [Training]](https://www.kaggle.com/ttahara/bms-mt-chemical-formula-regression-training)\n* [BMS-MT: Chemical Formula Regression [Inference]](https://www.kaggle.com/ttahara/bms-mt-chemical-formula-regression-inference)\n\nOOF Levenshtein distance for chemical formula is **0.9503** although I used **only 4%** of training data. The more data you use, the better score you can achieve.\n\nIn inference notebook, I tried solving the competition task by clustering approach using the predicted number of atoms as feature. This achieved slightly better (LB: 67.70) than other naive baselines.\n\n## What's next ?\n\nHere is local Levenshtein distance score of each part of `InChI` by regression and clustering approach:\n\n|     |   parts  |      mean_LS |\n|:---:|:--------:|:------------:|\n|0    | InChI_0  | 0.000000e+00 |\n|1    | InChI_1  | 9.503252e-01 |\n|2    | InChI_2  | **4.455307e+01** |\n|3    | InChI_3  | **1.762780e+01** |\n|4    | InChI_4  | 1.546733e+00 |\n|5    | InChI_5  | 3.841467e-01 |\n|6    | InChI_6  | 3.256986e-01 |\n|7    | InChI_7  | 2.150248e-02 |\n|8    | InChI_8  | 4.479854e-04 |\n|9    | InChI_9  | 3.011320e-05 |\n| 10  | InChI_10 | 8.250192e-07 |\n| 11  |   InChI  | 6.472638e+01 |\n\nAs you see, `InChI_2` (atom connections information)  and `InChI_3` (hydrogen atoms information) are the majority of the Levenshtein distance of whole the `InChI`. To focus on these, I'll try **multi-task approach**: multi-output regression for `InChI_1` and sequence prediction for `InChI_2` and `InChI_3`.\n\nI am very curious about which is better, multi-task approach or single sequence prediction approach  such as @yasufuminakama 's [baseline](https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-inference).",
    "1251771": "Finally a new kind of approach (it was getting boring as I couldn't think of any new); Thanks",
    "1252002": "I think deep understanding of `InChI` leads to various approaches other than single sequence predicton by encoder-decoder model.\n\nI hope there wil be a wide variety of solutions at the end of  competition :)",
    "1255352": "The input is an image, which can be represented as a vector. Theoretically, the output is a different representation of the same thing, so it should also be convertible to a vector. The trick is to convert InChI into a vector that can be converted back to InChI.",
    "1257210": "Are you saying some sort of Autoencoder can do it?",
    "1257212": "Are you saying some sort of Autoencoder can do it?",
    "1258180": "two separate ideas here:\n\n- think of each /c, /h, /m, etc as a query. \ngiven \"image + /x\", we are to retrive the structure for \"x\"\n\n- divide the substring of  InchI to words (i.e. entries of a dictionary). You can build a classifier to classifiy if the image contains the word or not. i.e. multi-class classification problem.\n\n- given a set of \"detected words\", predict the best arrangment\n\n---\n\ni wonder if anyone can suggest a good tokonizer for chemical InChI words?",
    "1258389": "hengck23\n- [WordPiece tokenizer for InChI](https://www.kaggle.com/affjljoo3581/bms-molecular-translation-train-inchi-tokenizer) kernel using AllenNLP\n- [BPE and others](https://pypi.org/project/tokenizers/)\n- [HuggingFace tokenizers](https://github.com/huggingface/tokenizers)",
    "1258633": "i wonder is there any specialised one for InCHI, e.g.\n\n![](https://github.com/XinhaoLi74/SmilesPE/raw/master/TOC.PNG)\nhttps://github.com/XinhaoLi74/SmilesPE"
  },
  "source": "meta"
}