{
  "id": 45728,
  "title": "Did anyone try text detection/classification approaches?",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/45728",
  "author_name": "",
  "post_date": "2017-12-15T02:02:57.064085600Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I found that \"MUSIQUE (CD)\" and \"LIBRAIRIE (books)\" categories had different nature from other categories. The accuracy in products of two categories was not good and probably they were needed problem specific processes. \nAlso, I found that @bestfitting posted text detection methods and word embedding libraries in external data thread (But we didn't use these approaches because of time lack).</p>\n\n<p>Did anyone try text detection/classification approaches?</p>",
  "messages": [
    {
      "id": "257855",
      "postDate": "12/15/2017 02:02:57",
      "content": "<p>I found that \"MUSIQUE (CD)\" and \"LIBRAIRIE (books)\" categories had different nature from other categories. The accuracy in products of two categories was not good and probably they were needed problem specific processes. \nAlso, I found that @bestfitting posted text detection methods and word embedding libraries in external data thread (But we didn't use these approaches because of time lack).</p>\n\n<p>Did anyone try text detection/classification approaches?</p>",
      "rawMarkdown": "I found that \"MUSIQUE (CD)\" and \"LIBRAIRIE (books)\" categories had different nature from other categories. The accuracy in products of two categories was not good and probably they were needed problem specific processes. \nAlso, I found that @bestfitting posted text detection methods and word embedding libraries in external data thread (But we didn't use these approaches because of time lack).\n\nDid anyone try text detection/classification approaches?",
      "votes": null
    },
    {
      "id": "257861",
      "postDate": "12/15/2017 02:13:58",
      "content": "<p>I tried Tesseract, but OCR itself didn't work well. Human could read a lot of text clearly, so I think it should be possible for NN to read them also. But I couldn't make it work reliably even after some preprocessings and gave it up.</p>",
      "rawMarkdown": "I tried Tesseract, but OCR itself didn't work well. Human could read a lot of text clearly, so I think it should be possible for NN to read them also. But I couldn't make it work reliably even after some preprocessings and gave it up.",
      "votes": null
    },
    {
      "id": "257868",
      "postDate": "12/15/2017 02:24:10",
      "content": "<p>I agree. It may work well if we train the NN properly.\nActually, Google Cloud Vision API can detect texts from Cdiscount images successfully.</p>",
      "rawMarkdown": "I agree. It may work well if we train the NN properly.\nActually, Google Cloud Vision API can detect texts from Cdiscount images successfully.",
      "votes": null
    },
    {
      "id": "257889",
      "postDate": "12/15/2017 03:01:51",
      "content": "<p>Although it took me a lot of time using OCR on CD and Book,it help me only with 0.001-0.0014 also.</p>",
      "rawMarkdown": "Although it took me a lot of time using OCR on CD and Book,it help me only with 0.001-0.0014 also.",
      "votes": null
    },
    {
      "id": "257895",
      "postDate": "12/15/2017 03:35:01",
      "content": "<p>Thank you for giving me your answer.\nWhy the score improvement was slight, I guess, maybe the majority of images in two categories are difficult to detect texts or nothing to be written.</p>\n\n<p>So that means your great results came from other factors, right? I'm glad if you will share your solution!</p>",
      "rawMarkdown": "Thank you for giving me your answer.\nWhy the score improvement was slight, I guess, maybe the majority of images in two categories are difficult to detect texts or nothing to be written.\n\nSo that means your great results came from other factors, right? I'm glad if you will share your solution!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 257861,
      "author_name": "jandjenter",
      "author_url": "",
      "post_date": "12/15/2017 02:13:58",
      "content": "<p>I tried Tesseract, but OCR itself didn't work well. Human could read a lot of text clearly, so I think it should be possible for NN to read them also. But I couldn't make it work reliably even after some preprocessings and gave it up.</p>",
      "votes": null,
      "replies": [
        {
          "id": 257868,
          "author_name": "lyakaap",
          "author_url": "",
          "post_date": "12/15/2017 02:24:10",
          "content": "<p>I agree. It may work well if we train the NN properly.\nActually, Google Cloud Vision API can detect texts from Cdiscount images successfully.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 257889,
      "author_name": "bestfitting",
      "author_url": "",
      "post_date": "12/15/2017 03:01:51",
      "content": "<p>Although it took me a lot of time using OCR on CD and Book,it help me only with 0.001-0.0014 also.</p>",
      "votes": null,
      "replies": [
        {
          "id": 257895,
          "author_name": "lyakaap",
          "author_url": "",
          "post_date": "12/15/2017 03:35:01",
          "content": "<p>Thank you for giving me your answer.\nWhy the score improvement was slight, I guess, maybe the majority of images in two categories are difficult to detect texts or nothing to be written.</p>\n\n<p>So that means your great results came from other factors, right? I'm glad if you will share your solution!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "257855": "I found that \"MUSIQUE (CD)\" and \"LIBRAIRIE (books)\" categories had different nature from other categories. The accuracy in products of two categories was not good and probably they were needed problem specific processes. \nAlso, I found that @bestfitting posted text detection methods and word embedding libraries in external data thread (But we didn't use these approaches because of time lack).\n\nDid anyone try text detection/classification approaches?",
    "257861": "I tried Tesseract, but OCR itself didn't work well. Human could read a lot of text clearly, so I think it should be possible for NN to read them also. But I couldn't make it work reliably even after some preprocessings and gave it up.",
    "257868": "I agree. It may work well if we train the NN properly.\nActually, Google Cloud Vision API can detect texts from Cdiscount images successfully.",
    "257889": "Although it took me a lot of time using OCR on CD and Book,it help me only with 0.001-0.0014 also.",
    "257895": "Thank you for giving me your answer.\nWhy the score improvement was slight, I guess, maybe the majority of images in two categories are difficult to detect texts or nothing to be written.\n\nSo that means your great results came from other factors, right? I'm glad if you will share your solution!"
  },
  "source": "meta"
}