{
  "id": 226216,
  "title": "Use of pytesseract (OCR) to get letters ?",
  "url": "/competitions/bms-molecular-translation/discussion/226216",
  "author_name": "",
  "post_date": "2021-03-15T15:32:46.073385Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Just a remark. </p>\n<p>Have someone tried to use pytesseract to get letters ? <br>\n(That is powerful OCR originally developped in 1980-90 and maintained and improved by google<br>\n<a href=\"https://en.wikipedia.org/wiki/Tesseract_(software\" target=\"_blank\">https://en.wikipedia.org/wiki/Tesseract_(software</a>) ). </p>\n<p>It is pre-installed on kaggle. <br>\nSo one can just:</p>\n<blockquote>\n  <p>import PIL<br>\n  import pytesseract<br>\n  pytesseract.image_to_string(image_sample) </p>\n</blockquote>\n<p>Some example for test: </p>\n<blockquote>\n  <p>import PIL<br>\n  import pytesseract<br>\n  import urllib<br>\n  urllib.request.urlretrieve(\"https://imgur.com/CfkJm1c.png\", \"sample2.png\")<br>\n  img = PIL.Image.open(\"sample2.png\")<br>\n  img.show()<br>\n  pytesseract.image_to_string(img) </p>\n</blockquote>\n<p>Will return: </p>\n<p>'import pytesseract\\n\\x0c'</p>",
  "messages": [
    {
      "id": "1239285",
      "postDate": "03/15/2021 15:32:46",
      "content": "<p>Just a remark. </p>\n<p>Have someone tried to use pytesseract to get letters ? <br>\n(That is powerful OCR originally developped in 1980-90 and maintained and improved by google<br>\n<a href=\"https://en.wikipedia.org/wiki/Tesseract_(software\" target=\"_blank\">https://en.wikipedia.org/wiki/Tesseract_(software</a>) ). </p>\n<p>It is pre-installed on kaggle. <br>\nSo one can just:</p>\n<blockquote>\n  <p>import PIL<br>\n  import pytesseract<br>\n  pytesseract.image_to_string(image_sample) </p>\n</blockquote>\n<p>Some example for test: </p>\n<blockquote>\n  <p>import PIL<br>\n  import pytesseract<br>\n  import urllib<br>\n  urllib.request.urlretrieve(\"https://imgur.com/CfkJm1c.png\", \"sample2.png\")<br>\n  img = PIL.Image.open(\"sample2.png\")<br>\n  img.show()<br>\n  pytesseract.image_to_string(img) </p>\n</blockquote>\n<p>Will return: </p>\n<p>'import pytesseract\\n\\x0c'</p>",
      "rawMarkdown": "Just a remark. \n\nHave someone tried to use pytesseract to get letters ? \n(That is powerful OCR originally developped in 1980-90 and maintained and improved by google\nhttps://en.wikipedia.org/wiki/Tesseract_(software) ). \n\nIt is pre-installed on kaggle. \nSo one can just:\n\n> import PIL\nimport pytesseract\npytesseract.image_to_string(image_sample) \n\n\nSome example for test: \n\n> import PIL\nimport pytesseract\nimport urllib\nurllib.request.urlretrieve(\"https://imgur.com/CfkJm1c.png\", \"sample2.png\")\nimg = PIL.Image.open(\"sample2.png\")\nimg.show()\npytesseract.image_to_string(img) \n\nWill return: \n\n'import pytesseract\\n\\x0c'",
      "votes": null
    },
    {
      "id": "1239291",
      "postDate": "03/15/2021 15:38:22",
      "content": "<p>I think that this is not a <strong>object character recognition</strong> task, because the image doesn't have any \"character\",  it has a graph that represents the <strong>InChI string</strong>. If someone thinks different please share.</p>",
      "rawMarkdown": "I think that this is not a **object character recognition** task, because the image doesn't have any \"character\",  it has a graph that represents the **InChI string**. If someone thinks different please share.",
      "votes": null
    },
    {
      "id": "1239310",
      "postDate": "03/15/2021 16:04:29",
      "content": "<p>I agree that the main task is to recognize the structure of the chemical formulas,<br>\nbut a kind of subtask (may be very small part of main task, but still ) we need to get the atom symbols C,H,O…<br>\nFor that part tessarct might be helpful. <br>\nAt least to check the obtained result. </p>",
      "rawMarkdown": "I agree that the main task is to recognize the structure of the chemical formulas,\nbut a kind of subtask (may be very small part of main task, but still ) we need to get the atom symbols C,H,O...\nFor that part tessarct might be helpful. \nAt least to check the obtained result.",
      "votes": null
    },
    {
      "id": "1264422",
      "postDate": "04/06/2021 06:25:58",
      "content": "<p>I always have trouble trying to get that package to work.</p>",
      "rawMarkdown": "I always have trouble trying to get that package to work.",
      "votes": null
    },
    {
      "id": "1264734",
      "postDate": "04/06/2021 12:01:15",
      "content": "<p>Do you plan to extract letters as features and use them to train the model? That's a good idea.<br>\nMy only concern is that the majority of models are already able to predict chemical formulas to a fairly high degree of accuracy.<br>\nI think the letters in the image are only effective for predicting the chemical formula, not the structure. So I am a little skeptical that it will improve accuracy.<br>\nHowever, I think your idea is interesting and informative form me. Thank you for sharing!</p>",
      "rawMarkdown": "Do you plan to extract letters as features and use them to train the model? That's a good idea.\nMy only concern is that the majority of models are already able to predict chemical formulas to a fairly high degree of accuracy.\nI think the letters in the image are only effective for predicting the chemical formula, not the structure. So I am a little skeptical that it will improve accuracy.\nHowever, I think your idea is interesting and informative form me. Thank you for sharing!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1239291,
      "author_name": "hiramcho",
      "author_url": "",
      "post_date": "03/15/2021 15:38:22",
      "content": "<p>I think that this is not a <strong>object character recognition</strong> task, because the image doesn't have any \"character\",  it has a graph that represents the <strong>InChI string</strong>. If someone thinks different please share.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1239310,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "03/15/2021 16:04:29",
          "content": "<p>I agree that the main task is to recognize the structure of the chemical formulas,<br>\nbut a kind of subtask (may be very small part of main task, but still ) we need to get the atom symbols C,H,O…<br>\nFor that part tessarct might be helpful. <br>\nAt least to check the obtained result. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1264422,
      "author_name": "eladwar",
      "author_url": "",
      "post_date": "04/06/2021 06:25:58",
      "content": "<p>I always have trouble trying to get that package to work.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1264734,
      "author_name": "yosukeyama",
      "author_url": "",
      "post_date": "04/06/2021 12:01:15",
      "content": "<p>Do you plan to extract letters as features and use them to train the model? That's a good idea.<br>\nMy only concern is that the majority of models are already able to predict chemical formulas to a fairly high degree of accuracy.<br>\nI think the letters in the image are only effective for predicting the chemical formula, not the structure. So I am a little skeptical that it will improve accuracy.<br>\nHowever, I think your idea is interesting and informative form me. Thank you for sharing!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1239285": "Just a remark. \n\nHave someone tried to use pytesseract to get letters ? \n(That is powerful OCR originally developped in 1980-90 and maintained and improved by google\nhttps://en.wikipedia.org/wiki/Tesseract_(software) ). \n\nIt is pre-installed on kaggle. \nSo one can just:\n\n> import PIL\nimport pytesseract\npytesseract.image_to_string(image_sample) \n\n\nSome example for test: \n\n> import PIL\nimport pytesseract\nimport urllib\nurllib.request.urlretrieve(\"https://imgur.com/CfkJm1c.png\", \"sample2.png\")\nimg = PIL.Image.open(\"sample2.png\")\nimg.show()\npytesseract.image_to_string(img) \n\nWill return: \n\n'import pytesseract\\n\\x0c'",
    "1239291": "I think that this is not a **object character recognition** task, because the image doesn't have any \"character\",  it has a graph that represents the **InChI string**. If someone thinks different please share.",
    "1239310": "I agree that the main task is to recognize the structure of the chemical formulas,\nbut a kind of subtask (may be very small part of main task, but still ) we need to get the atom symbols C,H,O...\nFor that part tessarct might be helpful. \nAt least to check the obtained result.",
    "1264422": "I always have trouble trying to get that package to work.",
    "1264734": "Do you plan to extract letters as features and use them to train the model? That's a good idea.\nMy only concern is that the majority of models are already able to predict chemical formulas to a fairly high degree of accuracy.\nI think the letters in the image are only effective for predicting the chemical formula, not the structure. So I am a little skeptical that it will improve accuracy.\nHowever, I think your idea is interesting and informative form me. Thank you for sharing!"
  },
  "source": "meta"
}