{
  "id": 105878,
  "title": "Presence of radicals in letters?",
  "url": "/competitions/kuzushiji-recognition/discussion/105878",
  "author_name": "",
  "post_date": "2019-08-27T01:35:02.334363Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Is is possible that there exist radicals in each different alphabet that is included in kuzushiji?</p>\n\n<p>For those who might not know, and might benefit from this information, a radical is a kind of \"primary\" shape on which later shapes can be drawn to create a character.</p>\n\n<p>Would find these patterns be a good strategy for this contest?</p>\n\n<p>Apologies for my ignorance. I don't read or speak japanese. Just admire ancient knowledge and would hate to see it lost</p>",
  "messages": [
    {
      "id": "608586",
      "postDate": "08/27/2019 01:35:02",
      "content": "<p>Is is possible that there exist radicals in each different alphabet that is included in kuzushiji?</p>\n\n<p>For those who might not know, and might benefit from this information, a radical is a kind of \"primary\" shape on which later shapes can be drawn to create a character.</p>\n\n<p>Would find these patterns be a good strategy for this contest?</p>\n\n<p>Apologies for my ignorance. I don't read or speak japanese. Just admire ancient knowledge and would hate to see it lost</p>",
      "rawMarkdown": "Is is possible that there exist radicals in each different alphabet that is included in kuzushiji?\n\nFor those who might not know, and might benefit from this information, a radical is a kind of \"primary\" shape on which later shapes can be drawn to create a character.\n\nWould find these patterns be a good strategy for this contest?\n\nApologies for my ignorance. I don't read or speak japanese. Just admire ancient knowledge and would hate to see it lost",
      "votes": null
    },
    {
      "id": "609793",
      "postDate": "08/28/2019 05:33:32",
      "content": "<p>I don't speak Japanese either but I did study Chinese for several years so I am familiar with many of characters. I actually had the same idea, to use radical/parts based hierarchal embedding for the characters. I haven't been able to find a comprehensive list, but it should be easy to build one via web scraping. I'll post it when I'm done. </p>",
      "rawMarkdown": "I don't speak Japanese either but I did study Chinese for several years so I am familiar with many of characters. I actually had the same idea, to use radical/parts based hierarchal embedding for the characters. I haven't been able to find a comprehensive list, but it should be easy to build one via web scraping. I'll post it when I'm done.",
      "votes": null
    },
    {
      "id": "614014",
      "postDate": "08/31/2019 02:39:37",
      "content": "<p>I've ended up using something called <a href=\"https://en.wikipedia.org/wiki/Chinese_character_description_language#Ideographic_Description_Sequences\">Ideographic Description Sequences</a> (IDS). A list of IDSs that includes all characters in the translation file can be found <a href=\"https://github.com/cjkvi/cjkvi-ids\">here</a>. </p>",
      "rawMarkdown": "I've ended up using something called [Ideographic Description Sequences](https://en.wikipedia.org/wiki/Chinese_character_description_language#Ideographic_Description_Sequences) (IDS). A list of IDSs that includes all characters in the translation file can be found [here](https://github.com/cjkvi/cjkvi-ids).",
      "votes": null
    },
    {
      "id": "632151",
      "postDate": "09/23/2019 09:33:27",
      "content": "<p>If you mean some stroke-level representation,  I don't think it makes much sense.</p>\n\n<p>There is little analogy between (e.g. Japanese/Chinese) radicals/strokes as to an alphabet and (e.g. English) characters as to a word. Japanese hiragana and katakana are phonetic, but both were simplified from Kanji. The order of strokes does matter in recognization task because that is in line with the direction in which text flows, but each radical has nothing to do with the meaning and beyond. It is like you seek to use some \"\\ + / + \\ + /\" for representing a \"w\" in a \"word\".</p>",
      "rawMarkdown": "If you mean some stroke-level representation,  I don't think it makes much sense.\n\nThere is little analogy between (e.g. Japanese/Chinese) radicals/strokes as to an alphabet and (e.g. English) characters as to a word. Japanese hiragana and katakana are phonetic, but both were simplified from Kanji. The order of strokes does matter in recognization task because that is in line with the direction in which text flows, but each radical has nothing to do with the meaning and beyond. It is like you seek to use some \"\\ + / + \\ + /\" for representing a \"w\" in a \"word\".",
      "votes": null
    },
    {
      "id": "641001",
      "postDate": "10/04/2019 11:25:04",
      "content": "<p>I think I had more or less the same idea.</p>\n\n<ol>\n<li><p>Detect individual characters.</p></li>\n<li><p>In each character detect the strokes (or the inflection points/corners with some well known technique such as <a href=\"https://docs.opencv.org/3.4/d4/d7d/tutorial_harris_detector.html\">https://docs.opencv.org/3.4/d4/d7d/tutorial_harris_detector.html</a>)</p></li>\n<li><p>Encode this data: each stroke/corner, its size, direction, angle..., and  distance to other stroke(s).  For hierarchical organization (which stroke \"depends\" on which other stroke...), use a simple algorithm, I saw such one on the \"molecular prediction\" competition of this summer.</p></li>\n<li><p>Feed the data to a net of some sort...</p></li>\n</ol>\n\n<p>Obviously this is far from a working implementation, and unfortunately I did not have time at all to work on this, and will not have time before end of competition (and the 30 hours GPU limit forces to be perfectly organized, even for a low-performing submission...)</p>\n\n<p>I would be curious to know if you pursued your approach and if it gave satisfying results.</p>",
      "rawMarkdown": "I think I had more or less the same idea.\n\n1. Detect individual characters.\n\n2. In each character detect the strokes (or the inflection points/corners with some well known technique such as https://docs.opencv.org/3.4/d4/d7d/tutorial_harris_detector.html)\n\n3. Encode this data: each stroke/corner, its size, direction, angle..., and  distance to other stroke(s).  For hierarchical organization (which stroke \"depends\" on which other stroke...), use a simple algorithm, I saw such one on the \"molecular prediction\" competition of this summer.\n\n4. Feed the data to a net of some sort...\n\nObviously this is far from a working implementation, and unfortunately I did not have time at all to work on this, and will not have time before end of competition (and the 30 hours GPU limit forces to be perfectly organized, even for a low-performing submission...)\n\nI would be curious to know if you pursued your approach and if it gave satisfying results.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 609793,
      "author_name": "deasmhumnha",
      "author_url": "",
      "post_date": "08/28/2019 05:33:32",
      "content": "<p>I don't speak Japanese either but I did study Chinese for several years so I am familiar with many of characters. I actually had the same idea, to use radical/parts based hierarchal embedding for the characters. I haven't been able to find a comprehensive list, but it should be easy to build one via web scraping. I'll post it when I'm done. </p>",
      "votes": null,
      "replies": [
        {
          "id": 641001,
          "author_name": "jtaglione",
          "author_url": "",
          "post_date": "10/04/2019 11:25:04",
          "content": "<p>I think I had more or less the same idea.</p>\n\n<ol>\n<li><p>Detect individual characters.</p></li>\n<li><p>In each character detect the strokes (or the inflection points/corners with some well known technique such as <a href=\"https://docs.opencv.org/3.4/d4/d7d/tutorial_harris_detector.html\">https://docs.opencv.org/3.4/d4/d7d/tutorial_harris_detector.html</a>)</p></li>\n<li><p>Encode this data: each stroke/corner, its size, direction, angle..., and  distance to other stroke(s).  For hierarchical organization (which stroke \"depends\" on which other stroke...), use a simple algorithm, I saw such one on the \"molecular prediction\" competition of this summer.</p></li>\n<li><p>Feed the data to a net of some sort...</p></li>\n</ol>\n\n<p>Obviously this is far from a working implementation, and unfortunately I did not have time at all to work on this, and will not have time before end of competition (and the 30 hours GPU limit forces to be perfectly organized, even for a low-performing submission...)</p>\n\n<p>I would be curious to know if you pursued your approach and if it gave satisfying results.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 614014,
      "author_name": "deasmhumnha",
      "author_url": "",
      "post_date": "08/31/2019 02:39:37",
      "content": "<p>I've ended up using something called <a href=\"https://en.wikipedia.org/wiki/Chinese_character_description_language#Ideographic_Description_Sequences\">Ideographic Description Sequences</a> (IDS). A list of IDSs that includes all characters in the translation file can be found <a href=\"https://github.com/cjkvi/cjkvi-ids\">here</a>. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 632151,
      "author_name": "anraninsg",
      "author_url": "",
      "post_date": "09/23/2019 09:33:27",
      "content": "<p>If you mean some stroke-level representation,  I don't think it makes much sense.</p>\n\n<p>There is little analogy between (e.g. Japanese/Chinese) radicals/strokes as to an alphabet and (e.g. English) characters as to a word. Japanese hiragana and katakana are phonetic, but both were simplified from Kanji. The order of strokes does matter in recognization task because that is in line with the direction in which text flows, but each radical has nothing to do with the meaning and beyond. It is like you seek to use some \"\\ + / + \\ + /\" for representing a \"w\" in a \"word\".</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "608586": "Is is possible that there exist radicals in each different alphabet that is included in kuzushiji?\n\nFor those who might not know, and might benefit from this information, a radical is a kind of \"primary\" shape on which later shapes can be drawn to create a character.\n\nWould find these patterns be a good strategy for this contest?\n\nApologies for my ignorance. I don't read or speak japanese. Just admire ancient knowledge and would hate to see it lost",
    "609793": "I don't speak Japanese either but I did study Chinese for several years so I am familiar with many of characters. I actually had the same idea, to use radical/parts based hierarchal embedding for the characters. I haven't been able to find a comprehensive list, but it should be easy to build one via web scraping. I'll post it when I'm done.",
    "614014": "I've ended up using something called [Ideographic Description Sequences](https://en.wikipedia.org/wiki/Chinese_character_description_language#Ideographic_Description_Sequences) (IDS). A list of IDSs that includes all characters in the translation file can be found [here](https://github.com/cjkvi/cjkvi-ids).",
    "632151": "If you mean some stroke-level representation,  I don't think it makes much sense.\n\nThere is little analogy between (e.g. Japanese/Chinese) radicals/strokes as to an alphabet and (e.g. English) characters as to a word. Japanese hiragana and katakana are phonetic, but both were simplified from Kanji. The order of strokes does matter in recognization task because that is in line with the direction in which text flows, but each radical has nothing to do with the meaning and beyond. It is like you seek to use some \"\\ + / + \\ + /\" for representing a \"w\" in a \"word\".",
    "641001": "I think I had more or less the same idea.\n\n1. Detect individual characters.\n\n2. In each character detect the strokes (or the inflection points/corners with some well known technique such as https://docs.opencv.org/3.4/d4/d7d/tutorial_harris_detector.html)\n\n3. Encode this data: each stroke/corner, its size, direction, angle..., and  distance to other stroke(s).  For hierarchical organization (which stroke \"depends\" on which other stroke...), use a simple algorithm, I saw such one on the \"molecular prediction\" competition of this summer.\n\n4. Feed the data to a net of some sort...\n\nObviously this is far from a working implementation, and unfortunately I did not have time at all to work on this, and will not have time before end of competition (and the 30 hours GPU limit forces to be perfectly organized, even for a low-performing submission...)\n\nI would be curious to know if you pursued your approach and if it gave satisfying results."
  },
  "source": "meta"
}