{
  "id": 100712,
  "title": "missing some unicode ids",
  "url": "/competitions/kuzushiji-recognition/discussion/100712",
  "author_name": "",
  "post_date": "2019-07-20T11:19:40.967762200Z",
  "votes": 11,
  "comment_count": 8,
  "views": 0,
  "content": "<p>In train.csv, I found [U+770C, U+4FA1, U+7A83, U+515A, U+5E81, U+5039], but I couldn't find these unicodes in unicode_translation.csv.</p>",
  "messages": [
    {
      "id": "580563",
      "postDate": "07/20/2019 11:19:40",
      "content": "<p>In train.csv, I found [U+770C, U+4FA1, U+7A83, U+515A, U+5E81, U+5039], but I couldn't find these unicodes in unicode_translation.csv.</p>",
      "rawMarkdown": "In train.csv, I found [U+770C, U+4FA1, U+7A83, U+515A, U+5E81, U+5039], but I couldn't find these unicodes in unicode_translation.csv.",
      "votes": null
    },
    {
      "id": "580634",
      "postDate": "07/20/2019 13:56:44",
      "content": "<p>Likely negligible - dictionary of Unicode Character and their corresponding frequency in the training lexicon <br> <br>[{'U+7A83': 1}, {'U+5039': 4}, {'U+770C': 10}, {'U+5E81': 3}, {'U+515A': 1}, {'U+4FA1': 3}]</p>",
      "rawMarkdown": "Likely negligible - dictionary of Unicode Character and their corresponding frequency in the training lexicon <br> <br>[{'U+7A83': 1}, {'U+5039': 4}, {'U+770C': 10}, {'U+5E81': 3}, {'U+515A': 1}, {'U+4FA1': 3}]",
      "votes": null
    },
    {
      "id": "580747",
      "postDate": "07/20/2019 17:19:07",
      "content": "<p>Thanks for bringing this to our attention, it should be fixed soon.</p>\n\n<p>In the meantime, you can convert any unicode codepoint such as <code>\"U+770C\"</code> to its corresponding character by using:\n<code>char = chr(int(unicode[2:], 16))</code></p>",
      "rawMarkdown": "Thanks for bringing this to our attention, it should be fixed soon.\n\nIn the meantime, you can convert any unicode codepoint such as `\"U+770C\"` to its corresponding character by using:\n```char = chr(int(unicode[2:], 16))```",
      "votes": null
    },
    {
      "id": "582107",
      "postDate": "07/22/2019 18:23:42",
      "content": "<p>I also see that few unicodes have the same \"char\" value, is it normal ?\nchar : 𩹵  Unicode  = ['U+24E30', 'U+2564A', 'U+28263', 'U+29E75']\nchar : 隆  Unicode = ['U+9686', 'U+F9DC']</p>",
      "rawMarkdown": "I also see that few unicodes have the same \"char\" value, is it normal ?\nchar : 𩹵  Unicode  = ['U+24E30', 'U+2564A', 'U+28263', 'U+29E75']\nchar : 隆  Unicode = ['U+9686', 'U+F9DC']",
      "votes": null
    },
    {
      "id": "583791",
      "postDate": "07/25/2019 01:44:33",
      "content": "<p>they're different according to (<a href=\"https://r12a.github.io/app-conversion/\">https://r12a.github.io/app-conversion/</a>)\nThat website says U+24E30 = 𤸰 while U+2564A = 𥙊</p>",
      "rawMarkdown": "they're different according to (https://r12a.github.io/app-conversion/)\nThat website says U+24E30 = 𤸰 while U+2564A = 𥙊",
      "votes": null
    },
    {
      "id": "583803",
      "postDate": "07/25/2019 02:28:14",
      "content": "<p>One problem with classical Japanese is many characters don't have unicode or even they do, modern fonts still can't show them correctly. They only show the closest one. Hence, they look the same in the text font. For example U+9686 and U+F9DC are different characters. If you look at font image closely (not the text font), there is a short line in the center of the character that make them different.  </p>",
      "rawMarkdown": "One problem with classical Japanese is many characters don't have unicode or even they do, modern fonts still can't show them correctly. They only show the closest one. Hence, they look the same in the text font. For example U+9686 and U+F9DC are different characters. If you look at font image closely (not the text font), there is a short line in the center of the character that make them different.",
      "votes": null
    },
    {
      "id": "584187",
      "postDate": "07/25/2019 14:35:12",
      "content": "<p>I'm uploading the corrected csv now, thank you for flagging this!</p>",
      "rawMarkdown": "I'm uploading the corrected csv now, thank you for flagging this!",
      "votes": null
    },
    {
      "id": "584405",
      "postDate": "07/25/2019 22:15:00",
      "content": "<p>I see, interesting, thanks ! </p>",
      "rawMarkdown": "I see, interesting, thanks !",
      "votes": null
    },
    {
      "id": "642679",
      "postDate": "10/06/2019 13:25:16",
      "content": "<p>When I run this : \n```\ntranslations_df = pd.read_csv('../input/kuzushiji-recognition/unicode_translation.csv')</p>\n\n<p>\"U+0031\" in str(translations_df[\"Unicode\"])</p>\n\n<p>\"U+4FA1\" in str(translations_df[\"Unicode\"])`\n```\nThe first one return True while the Second return False (same for U+770C). Is this normal ?</p>",
      "rawMarkdown": "When I run this : \n```\ntranslations_df = pd.read_csv('../input/kuzushiji-recognition/unicode_translation.csv')\n\n\"U+0031\" in str(translations_df[\"Unicode\"])\n\n\"U+4FA1\" in str(translations_df[\"Unicode\"])`\n```\nThe first one return True while the Second return False (same for U+770C). Is this normal ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 580634,
      "author_name": "teeyee314",
      "author_url": "",
      "post_date": "07/20/2019 13:56:44",
      "content": "<p>Likely negligible - dictionary of Unicode Character and their corresponding frequency in the training lexicon <br> <br>[{'U+7A83': 1}, {'U+5039': 4}, {'U+770C': 10}, {'U+5E81': 3}, {'U+515A': 1}, {'U+4FA1': 3}]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 580747,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "07/20/2019 17:19:07",
      "content": "<p>Thanks for bringing this to our attention, it should be fixed soon.</p>\n\n<p>In the meantime, you can convert any unicode codepoint such as <code>\"U+770C\"</code> to its corresponding character by using:\n<code>char = chr(int(unicode[2:], 16))</code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 582107,
      "author_name": "ludovick",
      "author_url": "",
      "post_date": "07/22/2019 18:23:42",
      "content": "<p>I also see that few unicodes have the same \"char\" value, is it normal ?\nchar : 𩹵  Unicode  = ['U+24E30', 'U+2564A', 'U+28263', 'U+29E75']\nchar : 隆  Unicode = ['U+9686', 'U+F9DC']</p>",
      "votes": null,
      "replies": [
        {
          "id": 583791,
          "author_name": "sidhanthholalkere",
          "author_url": "",
          "post_date": "07/25/2019 01:44:33",
          "content": "<p>they're different according to (<a href=\"https://r12a.github.io/app-conversion/\">https://r12a.github.io/app-conversion/</a>)\nThat website says U+24E30 = 𤸰 while U+2564A = 𥙊</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 583803,
          "author_name": "tkasasagi",
          "author_url": "",
          "post_date": "07/25/2019 02:28:14",
          "content": "<p>One problem with classical Japanese is many characters don't have unicode or even they do, modern fonts still can't show them correctly. They only show the closest one. Hence, they look the same in the text font. For example U+9686 and U+F9DC are different characters. If you look at font image closely (not the text font), there is a short line in the center of the character that make them different.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 584405,
          "author_name": "ludovick",
          "author_url": "",
          "post_date": "07/25/2019 22:15:00",
          "content": "<p>I see, interesting, thanks ! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 584187,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "07/25/2019 14:35:12",
      "content": "<p>I'm uploading the corrected csv now, thank you for flagging this!</p>",
      "votes": null,
      "replies": [
        {
          "id": 642679,
          "author_name": "cdk292",
          "author_url": "",
          "post_date": "10/06/2019 13:25:16",
          "content": "<p>When I run this : \n```\ntranslations_df = pd.read_csv('../input/kuzushiji-recognition/unicode_translation.csv')</p>\n\n<p>\"U+0031\" in str(translations_df[\"Unicode\"])</p>\n\n<p>\"U+4FA1\" in str(translations_df[\"Unicode\"])`\n```\nThe first one return True while the Second return False (same for U+770C). Is this normal ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "580563": "In train.csv, I found [U+770C, U+4FA1, U+7A83, U+515A, U+5E81, U+5039], but I couldn't find these unicodes in unicode_translation.csv.",
    "580634": "Likely negligible - dictionary of Unicode Character and their corresponding frequency in the training lexicon <br> <br>[{'U+7A83': 1}, {'U+5039': 4}, {'U+770C': 10}, {'U+5E81': 3}, {'U+515A': 1}, {'U+4FA1': 3}]",
    "580747": "Thanks for bringing this to our attention, it should be fixed soon.\n\nIn the meantime, you can convert any unicode codepoint such as `\"U+770C\"` to its corresponding character by using:\n```char = chr(int(unicode[2:], 16))```",
    "582107": "I also see that few unicodes have the same \"char\" value, is it normal ?\nchar : 𩹵  Unicode  = ['U+24E30', 'U+2564A', 'U+28263', 'U+29E75']\nchar : 隆  Unicode = ['U+9686', 'U+F9DC']",
    "583791": "they're different according to (https://r12a.github.io/app-conversion/)\nThat website says U+24E30 = 𤸰 while U+2564A = 𥙊",
    "583803": "One problem with classical Japanese is many characters don't have unicode or even they do, modern fonts still can't show them correctly. They only show the closest one. Hence, they look the same in the text font. For example U+9686 and U+F9DC are different characters. If you look at font image closely (not the text font), there is a short line in the center of the character that make them different.",
    "584187": "I'm uploading the corrected csv now, thank you for flagging this!",
    "584405": "I see, interesting, thanks !",
    "642679": "When I run this : \n```\ntranslations_df = pd.read_csv('../input/kuzushiji-recognition/unicode_translation.csv')\n\n\"U+0031\" in str(translations_df[\"Unicode\"])\n\n\"U+4FA1\" in str(translations_df[\"Unicode\"])`\n```\nThe first one return True while the Second return False (same for U+770C). Is this normal ?"
  },
  "source": "meta"
}