{
  "id": 103193,
  "title": "Ideas on character recognition?",
  "url": "/competitions/kuzushiji-recognition/discussion/103193",
  "author_name": "",
  "post_date": "2019-08-07T18:50:53.008470700Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I've never done a computer vision type problem before so I'm just curious if there are standard ways of finding the characters on the page? It seems like classifying them is pretty straightforward (similar to MNIST?) but I don't know about recognizing that a certain blob on the page is a legitimate character. </p>\n\n<p>Right now I'm using this tutorial <a href=\"https://pytorch.org/tutorials/intermediate/torchvision_tutorial.html\">https://pytorch.org/tutorials/intermediate/torchvision_tutorial.html</a> which uses a pre-trained network designed to detect pedestrians but I think it would work for letters too.</p>",
  "messages": [
    {
      "id": "594247",
      "postDate": "08/07/2019 18:50:53",
      "content": "<p>I've never done a computer vision type problem before so I'm just curious if there are standard ways of finding the characters on the page? It seems like classifying them is pretty straightforward (similar to MNIST?) but I don't know about recognizing that a certain blob on the page is a legitimate character. </p>\n\n<p>Right now I'm using this tutorial <a href=\"https://pytorch.org/tutorials/intermediate/torchvision_tutorial.html\">https://pytorch.org/tutorials/intermediate/torchvision_tutorial.html</a> which uses a pre-trained network designed to detect pedestrians but I think it would work for letters too.</p>",
      "rawMarkdown": "I've never done a computer vision type problem before so I'm just curious if there are standard ways of finding the characters on the page? It seems like classifying them is pretty straightforward (similar to MNIST?) but I don't know about recognizing that a certain blob on the page is a legitimate character. \n\nRight now I'm using this tutorial https://pytorch.org/tutorials/intermediate/torchvision_tutorial.html which uses a pre-trained network designed to detect pedestrians but I think it would work for letters too.",
      "votes": null
    },
    {
      "id": "594453",
      "postDate": "08/08/2019 03:48:16",
      "content": "<p>Is this a transfer learning approach? I think there's merit to the idea but I wonder how well the model will shift from recognizing pedestrians, which are bounded shapes regardless of their internal composition and patterns, to Kuzushiji characters which are more complex. Do let us know if you find success with it, and where you find it doesn't work as expected. </p>",
      "rawMarkdown": "Is this a transfer learning approach? I think there's merit to the idea but I wonder how well the model will shift from recognizing pedestrians, which are bounded shapes regardless of their internal composition and patterns, to Kuzushiji characters which are more complex. Do let us know if you find success with it, and where you find it doesn't work as expected.",
      "votes": null
    },
    {
      "id": "594689",
      "postDate": "08/08/2019 09:29:43",
      "content": "<p>Check out K_mat's \"CenterNet -Keypoint Detector\" Kernel. It's pretty good and shows you one possible solution</p>",
      "rawMarkdown": "Check out K_mat's \"CenterNet -Keypoint Detector\" Kernel. It's pretty good and shows you one possible solution",
      "votes": null
    },
    {
      "id": "595339",
      "postDate": "08/09/2019 05:47:34",
      "content": "<p><a href=\"/latimerb\">@latimerb</a>  This challenge involves us to locate and recognize the characters whereas in MNIST we just need to recognize the characters. You can use any object detection pretrained network for training on this dataset. you can use YOLO-v3 , SSD, Faster-rcnn for detection and recognition.Take a look at this post <a href=\"https://www.learnopencv.com/training-yolov3-deep-learning-based-custom-object-detector/\">https://www.learnopencv.com/training-yolov3-deep-learning-based-custom-object-detector/</a></p>",
      "rawMarkdown": "latimerb  This challenge involves us to locate and recognize the characters whereas in MNIST we just need to recognize the characters. You can use any object detection pretrained network for training on this dataset. you can use YOLO-v3 , SSD, Faster-rcnn for detection and recognition.Take a look at this post https://www.learnopencv.com/training-yolov3-deep-learning-based-custom-object-detector/",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 594453,
      "author_name": "jeremyloscheider",
      "author_url": "",
      "post_date": "08/08/2019 03:48:16",
      "content": "<p>Is this a transfer learning approach? I think there's merit to the idea but I wonder how well the model will shift from recognizing pedestrians, which are bounded shapes regardless of their internal composition and patterns, to Kuzushiji characters which are more complex. Do let us know if you find success with it, and where you find it doesn't work as expected. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 594689,
      "author_name": "christianwallenwein",
      "author_url": "",
      "post_date": "08/08/2019 09:29:43",
      "content": "<p>Check out K_mat's \"CenterNet -Keypoint Detector\" Kernel. It's pretty good and shows you one possible solution</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 595339,
      "author_name": "shivus",
      "author_url": "",
      "post_date": "08/09/2019 05:47:34",
      "content": "<p><a href=\"/latimerb\">@latimerb</a>  This challenge involves us to locate and recognize the characters whereas in MNIST we just need to recognize the characters. You can use any object detection pretrained network for training on this dataset. you can use YOLO-v3 , SSD, Faster-rcnn for detection and recognition.Take a look at this post <a href=\"https://www.learnopencv.com/training-yolov3-deep-learning-based-custom-object-detector/\">https://www.learnopencv.com/training-yolov3-deep-learning-based-custom-object-detector/</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "594247": "I've never done a computer vision type problem before so I'm just curious if there are standard ways of finding the characters on the page? It seems like classifying them is pretty straightforward (similar to MNIST?) but I don't know about recognizing that a certain blob on the page is a legitimate character. \n\nRight now I'm using this tutorial https://pytorch.org/tutorials/intermediate/torchvision_tutorial.html which uses a pre-trained network designed to detect pedestrians but I think it would work for letters too.",
    "594453": "Is this a transfer learning approach? I think there's merit to the idea but I wonder how well the model will shift from recognizing pedestrians, which are bounded shapes regardless of their internal composition and patterns, to Kuzushiji characters which are more complex. Do let us know if you find success with it, and where you find it doesn't work as expected.",
    "594689": "Check out K_mat's \"CenterNet -Keypoint Detector\" Kernel. It's pretty good and shows you one possible solution",
    "595339": "latimerb  This challenge involves us to locate and recognize the characters whereas in MNIST we just need to recognize the characters. You can use any object detection pretrained network for training on this dataset. you can use YOLO-v3 , SSD, Faster-rcnn for detection and recognition.Take a look at this post https://www.learnopencv.com/training-yolov3-deep-learning-based-custom-object-detector/"
  },
  "source": "meta"
}