{
  "id": 101488,
  "title": "Some Character Not Exist In Training Set",
  "url": "/competitions/kuzushiji-recognition/discussion/101488",
  "author_name": "",
  "post_date": "2019-07-26T08:34:57.974162300Z",
  "votes": 6,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I dont know if im doing this right, but i'm trying to check the frequency of the character from unicode_translation.csv in the training sets (train.csv). <br>\nI found that some of the unicode character dont appear in the training set, not even once.\nabout 575 character are not exist in the training set. how im supposed to classify it\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1023218%2Ff276aa56ab4623b0569fb05536072155%2FCapture-0%20Training%20Set.JPG?generation=1564130089854220&amp;alt=media\" alt=\"\">\nOther than that i saw that the number of the training sets are highly imbalanced\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1023218%2F986e96f9205470ddd260efbae061d116%2FCapture-1%20Training%20Set.JPG?generation=1564131047933334&amp;alt=media\" alt=\"\">\nsome characters like U+25CB appears 1236 times in the training set, while the other just appears once</p>",
  "messages": [
    {
      "id": "584625",
      "postDate": "07/26/2019 08:34:57",
      "content": "<p>I dont know if im doing this right, but i'm trying to check the frequency of the character from unicode_translation.csv in the training sets (train.csv). <br>\nI found that some of the unicode character dont appear in the training set, not even once.\nabout 575 character are not exist in the training set. how im supposed to classify it\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1023218%2Ff276aa56ab4623b0569fb05536072155%2FCapture-0%20Training%20Set.JPG?generation=1564130089854220&amp;alt=media\" alt=\"\">\nOther than that i saw that the number of the training sets are highly imbalanced\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1023218%2F986e96f9205470ddd260efbae061d116%2FCapture-1%20Training%20Set.JPG?generation=1564131047933334&amp;alt=media\" alt=\"\">\nsome characters like U+25CB appears 1236 times in the training set, while the other just appears once</p>",
      "rawMarkdown": "I dont know if im doing this right, but i'm trying to check the frequency of the character from unicode_translation.csv in the training sets (train.csv).  \nI found that some of the unicode character dont appear in the training set, not even once.\nabout 575 character are not exist in the training set. how im supposed to classify it\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1023218%2Ff276aa56ab4623b0569fb05536072155%2FCapture-0%20Training%20Set.JPG?generation=1564130089854220&amp;alt=media)\nOther than that i saw that the number of the training sets are highly imbalanced\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1023218%2F986e96f9205470ddd260efbae061d116%2FCapture-1%20Training%20Set.JPG?generation=1564131047933334&amp;alt=media)\nsome characters like U+25CB appears 1236 times in the training set, while the other just appears once",
      "votes": null
    },
    {
      "id": "585376",
      "postDate": "07/27/2019 10:36:55",
      "content": "<p>Only characters that are in the training set need to be classified and can appear in the test set. Not the entirety of the <code>unicode_transcription.csv</code>. You may safely ignore these additional characters</p>",
      "rawMarkdown": "Only characters that are in the training set need to be classified and can appear in the test set. Not the entirety of the `unicode_transcription.csv`. You may safely ignore these additional characters",
      "votes": null
    },
    {
      "id": "586275",
      "postDate": "07/28/2019 23:48:56",
      "content": "<p>thankyou</p>",
      "rawMarkdown": "thankyou",
      "votes": null
    },
    {
      "id": "594455",
      "postDate": "08/08/2019 03:49:05",
      "content": "<p>Out of curiosity, do you know the percentage of total characters that are represented by the characters in the training set?</p>",
      "rawMarkdown": "Out of curiosity, do you know the percentage of total characters that are represented by the characters in the training set?",
      "votes": null
    },
    {
      "id": "596463",
      "postDate": "08/10/2019 17:05:48",
      "content": "<p>Very close to 100%. The only characters that aren't in the training set are very rare Kanji that only appear once or twice in the entire dataset. Hence, ignoring these characters has very little effecton score, which is why we don't expect competitiors to find new character types that don't appear in the training set.</p>",
      "rawMarkdown": "Very close to 100%. The only characters that aren't in the training set are very rare Kanji that only appear once or twice in the entire dataset. Hence, ignoring these characters has very little effecton score, which is why we don't expect competitiors to find new character types that don't appear in the training set.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 585376,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "07/27/2019 10:36:55",
      "content": "<p>Only characters that are in the training set need to be classified and can appear in the test set. Not the entirety of the <code>unicode_transcription.csv</code>. You may safely ignore these additional characters</p>",
      "votes": null,
      "replies": [
        {
          "id": 594455,
          "author_name": "jeremyloscheider",
          "author_url": "",
          "post_date": "08/08/2019 03:49:05",
          "content": "<p>Out of curiosity, do you know the percentage of total characters that are represented by the characters in the training set?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 596463,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "08/10/2019 17:05:48",
          "content": "<p>Very close to 100%. The only characters that aren't in the training set are very rare Kanji that only appear once or twice in the entire dataset. Hence, ignoring these characters has very little effecton score, which is why we don't expect competitiors to find new character types that don't appear in the training set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 586275,
      "author_name": "arisjayandrana",
      "author_url": "",
      "post_date": "07/28/2019 23:48:56",
      "content": "<p>thankyou</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "584625": "I dont know if im doing this right, but i'm trying to check the frequency of the character from unicode_translation.csv in the training sets (train.csv).  \nI found that some of the unicode character dont appear in the training set, not even once.\nabout 575 character are not exist in the training set. how im supposed to classify it\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1023218%2Ff276aa56ab4623b0569fb05536072155%2FCapture-0%20Training%20Set.JPG?generation=1564130089854220&amp;alt=media)\nOther than that i saw that the number of the training sets are highly imbalanced\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1023218%2F986e96f9205470ddd260efbae061d116%2FCapture-1%20Training%20Set.JPG?generation=1564131047933334&amp;alt=media)\nsome characters like U+25CB appears 1236 times in the training set, while the other just appears once",
    "585376": "Only characters that are in the training set need to be classified and can appear in the test set. Not the entirety of the `unicode_transcription.csv`. You may safely ignore these additional characters",
    "586275": "thankyou",
    "594455": "Out of curiosity, do you know the percentage of total characters that are represented by the characters in the training set?",
    "596463": "Very close to 100%. The only characters that aren't in the training set are very rare Kanji that only appear once or twice in the entire dataset. Hence, ignoring these characters has very little effecton score, which is why we don't expect competitiors to find new character types that don't appear in the training set."
  },
  "source": "meta"
}