{
  "id": 240282,
  "title": "Found A Bug in YNakama's Tokenizer",
  "url": "/competitions/bms-molecular-translation/discussion/240282",
  "author_name": "",
  "post_date": "2021-05-19T06:46:21.167758500Z",
  "votes": 32,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I believe many (like us) started with YNakama's great notebook and used the tokenizer in it. So I guess it may be worth share that there exists a bug in the word segmentation part. </p>\n<p>For example, the training sample <code>002f57184f1f</code> has InChI <code>InChI=1S/C7H2Cl3F3N4/c8-6(9,10)5-16-3(7(11,12)13)2(1-14)4(15)17-5/h(H2,15,16,17)</code> and its word segmentation result should be</p>\n<pre><code>C 7 H 2 Cl 3 F 3 N 4 /c 8 - 6 ( 9 , 10 ) 5 - 16 - 3 ( 7 ( 11 , 12 ) 13 ) 2 ( 1 - 14 ) 4 ( 15 ) 17 - 5 /h ( H 2 , 15 , 16 , 17 )\n</code></pre>\n<p>while it actually is</p>\n<pre><code>C 7 H 2 Cl 3 F 3 N 4 /c 8 - 6 ( 9 , 10 ) 5 - 16 - 3 ( 7 ( 11 , 12 ) 13 ) 2 ( 1 - 14 ) 4 ( 15 ) 17 - 5 /h 2 , 15 , 16 , 17 )\n</code></pre>\n<p>.</p>\n<p>But I doubt this bug would have any real influences beacause it only affects 414 training samples and results in 790 Levenshtein distance in total. :P</p>",
  "messages": [
    {
      "id": "1314411",
      "postDate": "05/19/2021 06:46:21",
      "content": "<p>I believe many (like us) started with YNakama's great notebook and used the tokenizer in it. So I guess it may be worth share that there exists a bug in the word segmentation part. </p>\n<p>For example, the training sample <code>002f57184f1f</code> has InChI <code>InChI=1S/C7H2Cl3F3N4/c8-6(9,10)5-16-3(7(11,12)13)2(1-14)4(15)17-5/h(H2,15,16,17)</code> and its word segmentation result should be</p>\n<pre><code>C 7 H 2 Cl 3 F 3 N 4 /c 8 - 6 ( 9 , 10 ) 5 - 16 - 3 ( 7 ( 11 , 12 ) 13 ) 2 ( 1 - 14 ) 4 ( 15 ) 17 - 5 /h ( H 2 , 15 , 16 , 17 )\n</code></pre>\n<p>while it actually is</p>\n<pre><code>C 7 H 2 Cl 3 F 3 N 4 /c 8 - 6 ( 9 , 10 ) 5 - 16 - 3 ( 7 ( 11 , 12 ) 13 ) 2 ( 1 - 14 ) 4 ( 15 ) 17 - 5 /h 2 , 15 , 16 , 17 )\n</code></pre>\n<p>.</p>\n<p>But I doubt this bug would have any real influences beacause it only affects 414 training samples and results in 790 Levenshtein distance in total. :P</p>",
      "rawMarkdown": "I believe many (like us) started with YNakama's great notebook and used the tokenizer in it. So I guess it may be worth share that there exists a bug in the word segmentation part. \n\nFor example, the training sample `002f57184f1f` has InChI `InChI=1S/C7H2Cl3F3N4/c8-6(9,10)5-16-3(7(11,12)13)2(1-14)4(15)17-5/h(H2,15,16,17)` and its word segmentation result should be\n```\nC 7 H 2 Cl 3 F 3 N 4 /c 8 - 6 ( 9 , 10 ) 5 - 16 - 3 ( 7 ( 11 , 12 ) 13 ) 2 ( 1 - 14 ) 4 ( 15 ) 17 - 5 /h ( H 2 , 15 , 16 , 17 )\n```\nwhile it actually is\n```\nC 7 H 2 Cl 3 F 3 N 4 /c 8 - 6 ( 9 , 10 ) 5 - 16 - 3 ( 7 ( 11 , 12 ) 13 ) 2 ( 1 - 14 ) 4 ( 15 ) 17 - 5 /h 2 , 15 , 16 , 17 )\n```\n.\n\nBut I doubt this bug would have any real influences beacause it only affects 414 training samples and results in 790 Levenshtein distance in total. :P",
      "votes": null
    },
    {
      "id": "1314413",
      "postDate": "05/19/2021 06:48:52",
      "content": "<p>My code:</p>\n<pre><code>import re\nPATTEN = re.compile('\\d+|[A-Z][a-z]?|[^A-Za-z\\d/]|/[a-z]')\ndef l_split(s):\n    return ' '.join(re.findall(PATTEN,s))\n\ninchi = 'InChI=1S/C7H2Cl3F3N4/c8-6(9,10)5-16-3(7(11,12)13)2(1-14)4(15)17-5/h(H2,15,16,17)'\nsplited_inchi = l_split(inchi[9:]) \n\nsplited_inchi \n# 'C 7 H 2 Cl 3 F 3 N 4 /c 8 - 6 ( 9 , 10 ) 5 - 16 - 3 ( 7 ( 11 , 12 ) 13 ) 2 ( 1 - 14 ) 4 ( 15 ) 17 - 5 /h ( H 2 , 15 , 16 , 17 )'\n</code></pre>",
      "rawMarkdown": "My code:\n```ptyhon\nimport re\nPATTEN = re.compile('\\d+|[A-Z][a-z]?|[^A-Za-z\\d/]|/[a-z]')\ndef l_split(s):\n    return ' '.join(re.findall(PATTEN,s))\n\ninchi = 'InChI=1S/C7H2Cl3F3N4/c8-6(9,10)5-16-3(7(11,12)13)2(1-14)4(15)17-5/h(H2,15,16,17)'\nsplited_inchi = l_split(inchi[9:]) \n\nsplited_inchi \n# 'C 7 H 2 Cl 3 F 3 N 4 /c 8 - 6 ( 9 , 10 ) 5 - 16 - 3 ( 7 ( 11 , 12 ) 13 ) 2 ( 1 - 14 ) 4 ( 15 ) 17 - 5 /h ( H 2 , 15 , 16 , 17 )'\n```",
      "votes": null
    },
    {
      "id": "1314927",
      "postDate": "05/19/2021 13:08:21",
      "content": "<p>I didn't noticed it… thanks for sharing!</p>",
      "rawMarkdown": "I didn't noticed it... thanks for sharing!",
      "votes": null
    },
    {
      "id": "1315522",
      "postDate": "05/19/2021 21:32:33",
      "content": "<p>A good unit-test for (reversible) tokenization is to de-tokenize the tokenized text again and check if it matches the original:<br>\nInChI -&gt; tokenize -&gt; save -&gt; de-tokenize -&gt; InchI == De-tokenized-InchI?</p>\n<p>Comparable tests can be done for other common pre-processing steps.</p>",
      "rawMarkdown": "A good unit-test for (reversible) tokenization is to de-tokenize the tokenized text again and check if it matches the original:\nInChI -> tokenize -> save -> de-tokenize -> InchI == De-tokenized-InchI?\n\nComparable tests can be done for other common pre-processing steps.",
      "votes": null
    },
    {
      "id": "1315573",
      "postDate": "05/20/2021 00:02:15",
      "content": "<p>is there a fix for the bug?</p>",
      "rawMarkdown": "is there a fix for the bug?",
      "votes": null
    },
    {
      "id": "1315642",
      "postDate": "05/20/2021 03:24:34",
      "content": "<p>Maybe try the <code>l_split</code> function in the comment to split inchis? <br>\nI have not run the unit test as <a href=\"https://www.kaggle.com/cepheidq\" target=\"_blank\">@cepheidq</a> suggests yet, so I'm not one hundred percent sure.</p>",
      "rawMarkdown": "Maybe try the `l_split` function in the comment to split inchis? \nI have not run the unit test as @cepheidq suggests yet, so I'm not one hundred percent sure.",
      "votes": null
    },
    {
      "id": "1315644",
      "postDate": "05/20/2021 03:27:35",
      "content": "<p>Thank you for the great starter notebook. I learned a lot!😄</p>",
      "rawMarkdown": "Thank you for the great starter notebook. I learned a lot!😄",
      "votes": null
    },
    {
      "id": "1316169",
      "postDate": "05/20/2021 10:58:11",
      "content": "<p><a href=\"https://www.kaggle.com/ee1yii\" target=\"_blank\">@ee1yii</a> thanks a lot for sharing this! I did a quick check of your code and it looks like it provides the correct segmentation for those 414 molecules. I also doubt that fixing the tokenizer has any tangible effect on the score, but it is always good to make sure all the labels are correct :)</p>",
      "rawMarkdown": "ee1yii thanks a lot for sharing this! I did a quick check of your code and it looks like it provides the correct segmentation for those 414 molecules. I also doubt that fixing the tokenizer has any tangible effect on the score, but it is always good to make sure all the labels are correct :)",
      "votes": null
    },
    {
      "id": "1316257",
      "postDate": "05/20/2021 12:06:59",
      "content": "<p>I am using the YNakama's tokenizer, but I will not fix the bug and regard this as noise :) Seems to hurt not much because I am able to get 0.74 with a single model.</p>",
      "rawMarkdown": "I am using the YNakama's tokenizer, but I will not fix the bug and regard this as noise :) Seems to hurt not much because I am able to get 0.74 with a single model.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1314413,
      "author_name": "ee1yii",
      "author_url": "",
      "post_date": "05/19/2021 06:48:52",
      "content": "<p>My code:</p>\n<pre><code>import re\nPATTEN = re.compile('\\d+|[A-Z][a-z]?|[^A-Za-z\\d/]|/[a-z]')\ndef l_split(s):\n    return ' '.join(re.findall(PATTEN,s))\n\ninchi = 'InChI=1S/C7H2Cl3F3N4/c8-6(9,10)5-16-3(7(11,12)13)2(1-14)4(15)17-5/h(H2,15,16,17)'\nsplited_inchi = l_split(inchi[9:]) \n\nsplited_inchi \n# 'C 7 H 2 Cl 3 F 3 N 4 /c 8 - 6 ( 9 , 10 ) 5 - 16 - 3 ( 7 ( 11 , 12 ) 13 ) 2 ( 1 - 14 ) 4 ( 15 ) 17 - 5 /h ( H 2 , 15 , 16 , 17 )'\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1316169,
          "author_name": "kozodoi",
          "author_url": "",
          "post_date": "05/20/2021 10:58:11",
          "content": "<p><a href=\"https://www.kaggle.com/ee1yii\" target=\"_blank\">@ee1yii</a> thanks a lot for sharing this! I did a quick check of your code and it looks like it provides the correct segmentation for those 414 molecules. I also doubt that fixing the tokenizer has any tangible effect on the score, but it is always good to make sure all the labels are correct :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1314927,
      "author_name": "yasufuminakama",
      "author_url": "",
      "post_date": "05/19/2021 13:08:21",
      "content": "<p>I didn't noticed it… thanks for sharing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1315644,
          "author_name": "ee1yii",
          "author_url": "",
          "post_date": "05/20/2021 03:27:35",
          "content": "<p>Thank you for the great starter notebook. I learned a lot!😄</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1315522,
      "author_name": "cepheidq",
      "author_url": "",
      "post_date": "05/19/2021 21:32:33",
      "content": "<p>A good unit-test for (reversible) tokenization is to de-tokenize the tokenized text again and check if it matches the original:<br>\nInChI -&gt; tokenize -&gt; save -&gt; de-tokenize -&gt; InchI == De-tokenized-InchI?</p>\n<p>Comparable tests can be done for other common pre-processing steps.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1315573,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/20/2021 00:02:15",
      "content": "<p>is there a fix for the bug?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1315642,
          "author_name": "ee1yii",
          "author_url": "",
          "post_date": "05/20/2021 03:24:34",
          "content": "<p>Maybe try the <code>l_split</code> function in the comment to split inchis? <br>\nI have not run the unit test as <a href=\"https://www.kaggle.com/cepheidq\" target=\"_blank\">@cepheidq</a> suggests yet, so I'm not one hundred percent sure.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1316257,
      "author_name": "tugstugi",
      "author_url": "",
      "post_date": "05/20/2021 12:06:59",
      "content": "<p>I am using the YNakama's tokenizer, but I will not fix the bug and regard this as noise :) Seems to hurt not much because I am able to get 0.74 with a single model.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1314411": "I believe many (like us) started with YNakama's great notebook and used the tokenizer in it. So I guess it may be worth share that there exists a bug in the word segmentation part. \n\nFor example, the training sample `002f57184f1f` has InChI `InChI=1S/C7H2Cl3F3N4/c8-6(9,10)5-16-3(7(11,12)13)2(1-14)4(15)17-5/h(H2,15,16,17)` and its word segmentation result should be\n```\nC 7 H 2 Cl 3 F 3 N 4 /c 8 - 6 ( 9 , 10 ) 5 - 16 - 3 ( 7 ( 11 , 12 ) 13 ) 2 ( 1 - 14 ) 4 ( 15 ) 17 - 5 /h ( H 2 , 15 , 16 , 17 )\n```\nwhile it actually is\n```\nC 7 H 2 Cl 3 F 3 N 4 /c 8 - 6 ( 9 , 10 ) 5 - 16 - 3 ( 7 ( 11 , 12 ) 13 ) 2 ( 1 - 14 ) 4 ( 15 ) 17 - 5 /h 2 , 15 , 16 , 17 )\n```\n.\n\nBut I doubt this bug would have any real influences beacause it only affects 414 training samples and results in 790 Levenshtein distance in total. :P",
    "1314413": "My code:\n```ptyhon\nimport re\nPATTEN = re.compile('\\d+|[A-Z][a-z]?|[^A-Za-z\\d/]|/[a-z]')\ndef l_split(s):\n    return ' '.join(re.findall(PATTEN,s))\n\ninchi = 'InChI=1S/C7H2Cl3F3N4/c8-6(9,10)5-16-3(7(11,12)13)2(1-14)4(15)17-5/h(H2,15,16,17)'\nsplited_inchi = l_split(inchi[9:]) \n\nsplited_inchi \n# 'C 7 H 2 Cl 3 F 3 N 4 /c 8 - 6 ( 9 , 10 ) 5 - 16 - 3 ( 7 ( 11 , 12 ) 13 ) 2 ( 1 - 14 ) 4 ( 15 ) 17 - 5 /h ( H 2 , 15 , 16 , 17 )'\n```",
    "1314927": "I didn't noticed it... thanks for sharing!",
    "1315522": "A good unit-test for (reversible) tokenization is to de-tokenize the tokenized text again and check if it matches the original:\nInChI -> tokenize -> save -> de-tokenize -> InchI == De-tokenized-InchI?\n\nComparable tests can be done for other common pre-processing steps.",
    "1315573": "is there a fix for the bug?",
    "1315642": "Maybe try the `l_split` function in the comment to split inchis? \nI have not run the unit test as @cepheidq suggests yet, so I'm not one hundred percent sure.",
    "1315644": "Thank you for the great starter notebook. I learned a lot!😄",
    "1316169": "ee1yii thanks a lot for sharing this! I did a quick check of your code and it looks like it provides the correct segmentation for those 414 molecules. I also doubt that fixing the tokenizer has any tangible effect on the score, but it is always good to make sure all the labels are correct :)",
    "1316257": "I am using the YNakama's tokenizer, but I will not fix the bug and regard this as noise :) Seems to hurt not much because I am able to get 0.74 with a single model."
  },
  "source": "meta"
}