{
  "id": 103609,
  "title": "Ground truth annotation questions",
  "url": "/competitions/kuzushiji-recognition/discussion/103609",
  "author_name": "",
  "post_date": "2019-08-10T08:35:11.824995500Z",
  "votes": 11,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I'd like to clarify annotations on some pages in the dataset - are some of these annotation errors, or are there reasons why they were not annotated (or annotated), and if yes, could someone briefly clarify these? Ground truth boxes are in red, questions are in green. Thanks!</p>\n\n<p>Note that overall annotations look extremely clean, at least the bounding boxes - there are examples I found when analyzing model errors.</p>\n\n<p><code>200021763-00022_2</code>: two symbols in the middle of the text not annotated\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2Fbc58457940c4be84b9b6d94e9726e21a%2F200021763-00022_2.jpg?generation=1565425487003535&amp;alt=media\" alt=\"\"></p>\n\n<p><code>200021763-00023_1</code>: here as well some symbols in the middle of the text not annotated\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2Fc7dd145f6806aaf88e4df379bff3f064%2F200021763-00023_1.jpg?generation=1565425488751575&amp;alt=media\" alt=\"\"></p>\n\n<p><code>200015779_00087_2</code>: text in the \"comics\" not annotated:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F9a349262d8c5e343836da6cc23e1a86c%2F200015779_00087_2.jpg?generation=1565425488498504&amp;alt=media\" alt=\"\"></p>\n\n<p><code>200015779_00086_2</code>: some text in the image not annotated\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F77bc93bccbf16ab0bfaea12487cb9215%2F200015779_00086_2.jpg?generation=1565425487432015&amp;alt=media\" alt=\"\"></p>\n\n<p><code>100249537_00088_2</code>: some blocks of text not annotated\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2Ff6e4edd0cb80364b9fc266f66c38457c%2F100249537_00088_2.jpg?generation=1565425488482366&amp;alt=media\" alt=\"\"></p>\n\n<p><code>100249476_00027_2</code>: block of text at the bottom not annotated\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2Fe7660fef68800f7ad8a0976ec79eabf8%2F100249476_00027_2.jpg?generation=1565425488902470&amp;alt=media\" alt=\"\"></p>\n\n<p><code>200021644_00036_1</code>: some blocks of text not annotated\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F7fc5226429f274efed7539a9151a4a96%2F200021644_00036_1.jpg?generation=1565425490175621&amp;alt=media\" alt=\"\"></p>\n\n<p><code>200021644_00037_2</code>: some blocks of text not annotated\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F4f3efd59f7c2d0270abd7ec1ae1c4885%2F200021644_00037_2.jpg?generation=1565425489459771&amp;alt=media\" alt=\"\"></p>\n\n<p><code>200015779_00160_1</code>: probably not an error, but is this correct that some \"annotations\" to the right are annotated while some are not?\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F115ce0dc76b2a26f2845896b9221b1a5%2F200015779_00160_1.jpg?generation=1565425488749502&amp;alt=media\" alt=\"\"></p>\n\n<p><code>200015779_00047_2</code>: same question here\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F0d9653fa5dbb1009b73345e5fdd7dee4%2F200015779_00047_2.jpg?generation=1565425487977881&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "596168",
      "postDate": "08/10/2019 08:35:11",
      "content": "<p>I'd like to clarify annotations on some pages in the dataset - are some of these annotation errors, or are there reasons why they were not annotated (or annotated), and if yes, could someone briefly clarify these? Ground truth boxes are in red, questions are in green. Thanks!</p>\n\n<p>Note that overall annotations look extremely clean, at least the bounding boxes - there are examples I found when analyzing model errors.</p>\n\n<p><code>200021763-00022_2</code>: two symbols in the middle of the text not annotated\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2Fbc58457940c4be84b9b6d94e9726e21a%2F200021763-00022_2.jpg?generation=1565425487003535&amp;alt=media\" alt=\"\"></p>\n\n<p><code>200021763-00023_1</code>: here as well some symbols in the middle of the text not annotated\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2Fc7dd145f6806aaf88e4df379bff3f064%2F200021763-00023_1.jpg?generation=1565425488751575&amp;alt=media\" alt=\"\"></p>\n\n<p><code>200015779_00087_2</code>: text in the \"comics\" not annotated:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F9a349262d8c5e343836da6cc23e1a86c%2F200015779_00087_2.jpg?generation=1565425488498504&amp;alt=media\" alt=\"\"></p>\n\n<p><code>200015779_00086_2</code>: some text in the image not annotated\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F77bc93bccbf16ab0bfaea12487cb9215%2F200015779_00086_2.jpg?generation=1565425487432015&amp;alt=media\" alt=\"\"></p>\n\n<p><code>100249537_00088_2</code>: some blocks of text not annotated\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2Ff6e4edd0cb80364b9fc266f66c38457c%2F100249537_00088_2.jpg?generation=1565425488482366&amp;alt=media\" alt=\"\"></p>\n\n<p><code>100249476_00027_2</code>: block of text at the bottom not annotated\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2Fe7660fef68800f7ad8a0976ec79eabf8%2F100249476_00027_2.jpg?generation=1565425488902470&amp;alt=media\" alt=\"\"></p>\n\n<p><code>200021644_00036_1</code>: some blocks of text not annotated\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F7fc5226429f274efed7539a9151a4a96%2F200021644_00036_1.jpg?generation=1565425490175621&amp;alt=media\" alt=\"\"></p>\n\n<p><code>200021644_00037_2</code>: some blocks of text not annotated\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F4f3efd59f7c2d0270abd7ec1ae1c4885%2F200021644_00037_2.jpg?generation=1565425489459771&amp;alt=media\" alt=\"\"></p>\n\n<p><code>200015779_00160_1</code>: probably not an error, but is this correct that some \"annotations\" to the right are annotated while some are not?\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F115ce0dc76b2a26f2845896b9221b1a5%2F200015779_00160_1.jpg?generation=1565425488749502&amp;alt=media\" alt=\"\"></p>\n\n<p><code>200015779_00047_2</code>: same question here\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F0d9653fa5dbb1009b73345e5fdd7dee4%2F200015779_00047_2.jpg?generation=1565425487977881&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I'd like to clarify annotations on some pages in the dataset - are some of these annotation errors, or are there reasons why they were not annotated (or annotated), and if yes, could someone briefly clarify these? Ground truth boxes are in red, questions are in green. Thanks!\n\nNote that overall annotations look extremely clean, at least the bounding boxes - there are examples I found when analyzing model errors.\n\n``200021763-00022_2``: two symbols in the middle of the text not annotated\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2Fbc58457940c4be84b9b6d94e9726e21a%2F200021763-00022_2.jpg?generation=1565425487003535&amp;alt=media)\n\n``200021763-00023_1``: here as well some symbols in the middle of the text not annotated\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2Fc7dd145f6806aaf88e4df379bff3f064%2F200021763-00023_1.jpg?generation=1565425488751575&amp;alt=media)\n\n``200015779_00087_2``: text in the \"comics\" not annotated:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F9a349262d8c5e343836da6cc23e1a86c%2F200015779_00087_2.jpg?generation=1565425488498504&amp;alt=media)\n\n``200015779_00086_2``: some text in the image not annotated\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F77bc93bccbf16ab0bfaea12487cb9215%2F200015779_00086_2.jpg?generation=1565425487432015&amp;alt=media)\n\n``100249537_00088_2``: some blocks of text not annotated\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2Ff6e4edd0cb80364b9fc266f66c38457c%2F100249537_00088_2.jpg?generation=1565425488482366&amp;alt=media)\n\n``100249476_00027_2``: block of text at the bottom not annotated\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2Fe7660fef68800f7ad8a0976ec79eabf8%2F100249476_00027_2.jpg?generation=1565425488902470&amp;alt=media)\n\n``200021644_00036_1``: some blocks of text not annotated\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F7fc5226429f274efed7539a9151a4a96%2F200021644_00036_1.jpg?generation=1565425490175621&amp;alt=media)\n\n``200021644_00037_2``: some blocks of text not annotated\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F4f3efd59f7c2d0270abd7ec1ae1c4885%2F200021644_00037_2.jpg?generation=1565425489459771&amp;alt=media)\n\n``200015779_00160_1``: probably not an error, but is this correct that some \"annotations\" to the right are annotated while some are not?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F115ce0dc76b2a26f2845896b9221b1a5%2F200015779_00160_1.jpg?generation=1565425488749502&amp;alt=media)\n\n``200015779_00047_2``: same question here\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F0d9653fa5dbb1009b73345e5fdd7dee4%2F200015779_00047_2.jpg?generation=1565425487977881&amp;alt=media)",
      "votes": null
    },
    {
      "id": "596479",
      "postDate": "08/10/2019 17:34:10",
      "content": "<p>Let's hope Competition organisers like <a href=\"/anokas\">@anokas</a> will clarrify about this soon</p>",
      "rawMarkdown": "Let's hope Competition organisers like @anokas will clarrify about this soon",
      "votes": null
    },
    {
      "id": "596563",
      "postDate": "08/10/2019 21:18:11",
      "content": "<p>I have the same question.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2938236%2F1674e2b2d66391ef0028e1852a142174%2Fb.png?generation=1565471860491867&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I have the same question.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2938236%2F1674e2b2d66391ef0028e1852a142174%2Fb.png?generation=1565471860491867&amp;alt=media)",
      "votes": null
    },
    {
      "id": "596964",
      "postDate": "08/11/2019 15:36:33",
      "content": "<p>As the end goal is to translate and digitise ancient books, I think there is a particular focus on the main columns of symbols while the figures and notes are ignored. Regarding the non-annotated symbols within the main text, my wild guess is that they are family or place names, or some types of kuzushiji symbols that were used so little that no one know their meanings anymore. \nIf I understood well <a href=\"https://twitter.com/tkasasagi/status/1141569771546136576\">this post</a> by <a href=\"/tkasasagi\">@tkasasagi</a> explaining some details about the manual annotation process, the smaller symbols on the right of a kuzushiji symbol may be a note in modern Japanese. However, if that is the case, I have no idea why some are annotated and some aren't. </p>",
      "rawMarkdown": "As the end goal is to translate and digitise ancient books, I think there is a particular focus on the main columns of symbols while the figures and notes are ignored. Regarding the non-annotated symbols within the main text, my wild guess is that they are family or place names, or some types of kuzushiji symbols that were used so little that no one know their meanings anymore. \nIf I understood well [this post](https://twitter.com/tkasasagi/status/1141569771546136576) by @tkasasagi explaining some details about the manual annotation process, the smaller symbols on the right of a kuzushiji symbol may be a note in modern Japanese. However, if that is the case, I have no idea why some are annotated and some aren't.",
      "votes": null
    },
    {
      "id": "597608",
      "postDate": "08/12/2019 15:02:34",
      "content": "<p>I can only speak to a couple of cases:\n- Some characters were sufficiently rare that they were dropped from the labels, regardless of location. \n- Per the data page, annotations outside of main writing grids were intentionally left out.   </p>",
      "rawMarkdown": "I can only speak to a couple of cases:\n- Some characters were sufficiently rare that they were dropped from the labels, regardless of location. \n- Per the data page, annotations outside of main writing grids were intentionally left out.",
      "votes": null
    },
    {
      "id": "606319",
      "postDate": "08/23/2019 12:30:50",
      "content": "<p>There are a few reasons for this. Some characters those are not in GT mostly because not every character in Classical Japanese has Unicode. Like Chinese, there are a lot of characters in Japanese too. Many characters are not used in modern Japanese anymore so there is no code for them. Since there is no proper character we can put it in the GT for them so we dropped them. </p>\n\n<p>Another reason is because when the dataset was created, we only considered main text as a part of the dataset. For example, 100249537_00088_2 and 100249476_00027_2 have handwritten annotation (reader annotations) on the top block so we don't have GT for it. </p>\n\n<p>Another type of missing GT is description for illustrations which is also not a part of main text. Like in 200021644_00036_1, 200021644_00037_2 which the text is a part of illustration not actual text of the book.</p>\n\n<p>For 200015779_00160_1 is really hard to judge whether it's annotation or not. Not all small characters are annotation. One thing experts use to judge is whether it has some text on the left side or not. Sometime, the characters were added later, but if it's actually part of the main text so there is GT for it. </p>\n\n<p>Historical documents are very dynamic and not everything can be mapped to modern language.</p>\n\n<p>For the one that <a href=\"/kmat2019\">@kmat2019</a> has question. I think it's completely mislabel. </p>",
      "rawMarkdown": "There are a few reasons for this. Some characters those are not in GT mostly because not every character in Classical Japanese has Unicode. Like Chinese, there are a lot of characters in Japanese too. Many characters are not used in modern Japanese anymore so there is no code for them. Since there is no proper character we can put it in the GT for them so we dropped them. \n\nAnother reason is because when the dataset was created, we only considered main text as a part of the dataset. For example, 100249537_00088_2 and 100249476_00027_2 have handwritten annotation (reader annotations) on the top block so we don't have GT for it. \n\nAnother type of missing GT is description for illustrations which is also not a part of main text. Like in 200021644_00036_1, 200021644_00037_2 which the text is a part of illustration not actual text of the book.\n\nFor 200015779_00160_1 is really hard to judge whether it's annotation or not. Not all small characters are annotation. One thing experts use to judge is whether it has some text on the left side or not. Sometime, the characters were added later, but if it's actually part of the main text so there is GT for it. \n\nHistorical documents are very dynamic and not everything can be mapped to modern language.\n\nFor the one that @kmat2019 has question. I think it's completely mislabel.",
      "votes": null
    },
    {
      "id": "606323",
      "postDate": "08/23/2019 12:41:55",
      "content": "<p>Makes sense, thank you so much for a very detailed reply.</p>",
      "rawMarkdown": "Makes sense, thank you so much for a very detailed reply.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 596479,
      "author_name": "kurianbenoy",
      "author_url": "",
      "post_date": "08/10/2019 17:34:10",
      "content": "<p>Let's hope Competition organisers like <a href=\"/anokas\">@anokas</a> will clarrify about this soon</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 596563,
      "author_name": "kmat2019",
      "author_url": "",
      "post_date": "08/10/2019 21:18:11",
      "content": "<p>I have the same question.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2938236%2F1674e2b2d66391ef0028e1852a142174%2Fb.png?generation=1565471860491867&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 596964,
      "author_name": "frlemarchand",
      "author_url": "",
      "post_date": "08/11/2019 15:36:33",
      "content": "<p>As the end goal is to translate and digitise ancient books, I think there is a particular focus on the main columns of symbols while the figures and notes are ignored. Regarding the non-annotated symbols within the main text, my wild guess is that they are family or place names, or some types of kuzushiji symbols that were used so little that no one know their meanings anymore. \nIf I understood well <a href=\"https://twitter.com/tkasasagi/status/1141569771546136576\">this post</a> by <a href=\"/tkasasagi\">@tkasasagi</a> explaining some details about the manual annotation process, the smaller symbols on the right of a kuzushiji symbol may be a note in modern Japanese. However, if that is the case, I have no idea why some are annotated and some aren't. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 597608,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "08/12/2019 15:02:34",
      "content": "<p>I can only speak to a couple of cases:\n- Some characters were sufficiently rare that they were dropped from the labels, regardless of location. \n- Per the data page, annotations outside of main writing grids were intentionally left out.   </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 606319,
      "author_name": "tkasasagi",
      "author_url": "",
      "post_date": "08/23/2019 12:30:50",
      "content": "<p>There are a few reasons for this. Some characters those are not in GT mostly because not every character in Classical Japanese has Unicode. Like Chinese, there are a lot of characters in Japanese too. Many characters are not used in modern Japanese anymore so there is no code for them. Since there is no proper character we can put it in the GT for them so we dropped them. </p>\n\n<p>Another reason is because when the dataset was created, we only considered main text as a part of the dataset. For example, 100249537_00088_2 and 100249476_00027_2 have handwritten annotation (reader annotations) on the top block so we don't have GT for it. </p>\n\n<p>Another type of missing GT is description for illustrations which is also not a part of main text. Like in 200021644_00036_1, 200021644_00037_2 which the text is a part of illustration not actual text of the book.</p>\n\n<p>For 200015779_00160_1 is really hard to judge whether it's annotation or not. Not all small characters are annotation. One thing experts use to judge is whether it has some text on the left side or not. Sometime, the characters were added later, but if it's actually part of the main text so there is GT for it. </p>\n\n<p>Historical documents are very dynamic and not everything can be mapped to modern language.</p>\n\n<p>For the one that <a href=\"/kmat2019\">@kmat2019</a> has question. I think it's completely mislabel. </p>",
      "votes": null,
      "replies": [
        {
          "id": 606323,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "08/23/2019 12:41:55",
          "content": "<p>Makes sense, thank you so much for a very detailed reply.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "596168": "I'd like to clarify annotations on some pages in the dataset - are some of these annotation errors, or are there reasons why they were not annotated (or annotated), and if yes, could someone briefly clarify these? Ground truth boxes are in red, questions are in green. Thanks!\n\nNote that overall annotations look extremely clean, at least the bounding boxes - there are examples I found when analyzing model errors.\n\n``200021763-00022_2``: two symbols in the middle of the text not annotated\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2Fbc58457940c4be84b9b6d94e9726e21a%2F200021763-00022_2.jpg?generation=1565425487003535&amp;alt=media)\n\n``200021763-00023_1``: here as well some symbols in the middle of the text not annotated\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2Fc7dd145f6806aaf88e4df379bff3f064%2F200021763-00023_1.jpg?generation=1565425488751575&amp;alt=media)\n\n``200015779_00087_2``: text in the \"comics\" not annotated:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F9a349262d8c5e343836da6cc23e1a86c%2F200015779_00087_2.jpg?generation=1565425488498504&amp;alt=media)\n\n``200015779_00086_2``: some text in the image not annotated\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F77bc93bccbf16ab0bfaea12487cb9215%2F200015779_00086_2.jpg?generation=1565425487432015&amp;alt=media)\n\n``100249537_00088_2``: some blocks of text not annotated\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2Ff6e4edd0cb80364b9fc266f66c38457c%2F100249537_00088_2.jpg?generation=1565425488482366&amp;alt=media)\n\n``100249476_00027_2``: block of text at the bottom not annotated\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2Fe7660fef68800f7ad8a0976ec79eabf8%2F100249476_00027_2.jpg?generation=1565425488902470&amp;alt=media)\n\n``200021644_00036_1``: some blocks of text not annotated\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F7fc5226429f274efed7539a9151a4a96%2F200021644_00036_1.jpg?generation=1565425490175621&amp;alt=media)\n\n``200021644_00037_2``: some blocks of text not annotated\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F4f3efd59f7c2d0270abd7ec1ae1c4885%2F200021644_00037_2.jpg?generation=1565425489459771&amp;alt=media)\n\n``200015779_00160_1``: probably not an error, but is this correct that some \"annotations\" to the right are annotated while some are not?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F115ce0dc76b2a26f2845896b9221b1a5%2F200015779_00160_1.jpg?generation=1565425488749502&amp;alt=media)\n\n``200015779_00047_2``: same question here\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F0d9653fa5dbb1009b73345e5fdd7dee4%2F200015779_00047_2.jpg?generation=1565425487977881&amp;alt=media)",
    "596479": "Let's hope Competition organisers like @anokas will clarrify about this soon",
    "596563": "I have the same question.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2938236%2F1674e2b2d66391ef0028e1852a142174%2Fb.png?generation=1565471860491867&amp;alt=media)",
    "596964": "As the end goal is to translate and digitise ancient books, I think there is a particular focus on the main columns of symbols while the figures and notes are ignored. Regarding the non-annotated symbols within the main text, my wild guess is that they are family or place names, or some types of kuzushiji symbols that were used so little that no one know their meanings anymore. \nIf I understood well [this post](https://twitter.com/tkasasagi/status/1141569771546136576) by @tkasasagi explaining some details about the manual annotation process, the smaller symbols on the right of a kuzushiji symbol may be a note in modern Japanese. However, if that is the case, I have no idea why some are annotated and some aren't.",
    "597608": "I can only speak to a couple of cases:\n- Some characters were sufficiently rare that they were dropped from the labels, regardless of location. \n- Per the data page, annotations outside of main writing grids were intentionally left out.",
    "606319": "There are a few reasons for this. Some characters those are not in GT mostly because not every character in Classical Japanese has Unicode. Like Chinese, there are a lot of characters in Japanese too. Many characters are not used in modern Japanese anymore so there is no code for them. Since there is no proper character we can put it in the GT for them so we dropped them. \n\nAnother reason is because when the dataset was created, we only considered main text as a part of the dataset. For example, 100249537_00088_2 and 100249476_00027_2 have handwritten annotation (reader annotations) on the top block so we don't have GT for it. \n\nAnother type of missing GT is description for illustrations which is also not a part of main text. Like in 200021644_00036_1, 200021644_00037_2 which the text is a part of illustration not actual text of the book.\n\nFor 200015779_00160_1 is really hard to judge whether it's annotation or not. Not all small characters are annotation. One thing experts use to judge is whether it has some text on the left side or not. Sometime, the characters were added later, but if it's actually part of the main text so there is GT for it. \n\nHistorical documents are very dynamic and not everything can be mapped to modern language.\n\nFor the one that @kmat2019 has question. I think it's completely mislabel.",
    "606323": "Makes sense, thank you so much for a very detailed reply."
  },
  "source": "meta"
}