{
  "id": 106103,
  "title": "Using a language model",
  "url": "/competitions/kuzushiji-recognition/discussion/106103",
  "author_name": "Konstantin Lopukhin",
  "post_date": "2019-08-28T07:34:21.760000",
  "votes": 18,
  "comment_count": 22,
  "views": 0,
  "content": "<p>I wonder if anyone was able to successfully use a language model? It should be helpful because as I understand, in many cases a character would be easier to recognize knowing it's context, and this is a technique commonly used with OCR, see for example <a href=\"https://medium.com/apache-mxnet/handwriting-ocr-handwriting-recognition-and-language-modeling-with-mxnet-gluon-4c7165788c67\">https://medium.com/apache-mxnet/handwriting-ocr-handwriting-recognition-and-language-modeling-with-mxnet-gluon-4c7165788c67</a>\nI tried building a simple LSTM language model, but the best validation loss (cross-entropy) I could get was 4.8 (with a per-book validation split), which sounds extremely bad compared to 0.6402 cross-entropy loss I get with the same split with the image model (which gets my current best LB score of 0.863, it's a single fold single model).</p>\n\n<p>I'm interested in your thoughts on the following:\n- do you think a language model should help here?\n- did you try building a language model and if yes, what kind of log loss / perplexity did you get?\n- did you try incorporating the language model into your solution, and if yes, did it improve the score?</p>",
  "messages": [
    {
      "id": 609876,
      "postDate": "2019-08-28T07:34:21.760Z",
      "content": "<p>I wonder if anyone was able to successfully use a language model? It should be helpful because as I understand, in many cases a character would be easier to recognize knowing it's context, and this is a technique commonly used with OCR, see for example <a href=\"https://medium.com/apache-mxnet/handwriting-ocr-handwriting-recognition-and-language-modeling-with-mxnet-gluon-4c7165788c67\">https://medium.com/apache-mxnet/handwriting-ocr-handwriting-recognition-and-language-modeling-with-mxnet-gluon-4c7165788c67</a>\nI tried building a simple LSTM language model, but the best validation loss (cross-entropy) I could get was 4.8 (with a per-book validation split), which sounds extremely bad compared to 0.6402 cross-entropy loss I get with the same split with the image model (which gets my current best LB score of 0.863, it's a single fold single model).</p>\n\n<p>I'm interested in your thoughts on the following:\n- do you think a language model should help here?\n- did you try building a language model and if yes, what kind of log loss / perplexity did you get?\n- did you try incorporating the language model into your solution, and if yes, did it improve the score?</p>",
      "rawMarkdown": "I wonder if anyone was able to successfully use a language model? It should be helpful because as I understand, in many cases a character would be easier to recognize knowing it's context, and this is a technique commonly used with OCR, see for example https://medium.com/apache-mxnet/handwriting-ocr-handwriting-recognition-and-language-modeling-with-mxnet-gluon-4c7165788c67\nI tried building a simple LSTM language model, but the best validation loss (cross-entropy) I could get was 4.8 (with a per-book validation split), which sounds extremely bad compared to 0.6402 cross-entropy loss I get with the same split with the image model (which gets my current best LB score of 0.863, it's a single fold single model).\n\nI'm interested in your thoughts on the following:\n- do you think a language model should help here?\n- did you try building a language model and if yes, what kind of log loss / perplexity did you get?\n- did you try incorporating the language model into your solution, and if yes, did it improve the score?",
      "votes": 17
    },
    {
      "id": 610096,
      "postDate": "2019-08-28T12:15:00.047Z",
      "content": "<blockquote>\n  <p>do you think a language model should help here?</p>\n</blockquote>\n\n<p>Yes, I think so. but I've not built a model yet.\nLet me explain with the first sentence as follows. </p>\n\n<p><code>[自序] [若]い[時]の[気強]に[己]やれと[思]ふた[細工]も[老武者]の\nかなしさは[息子]に[及]ず[浮世]を[裏]の[三畳]に[避]て[正風]の\n[俳諧]を[楽]しめども[根]が[職人]の[文盲]だけこそけれ</code></p>\n\n<p>The kanji are written in parentheses, the others are kana (common easy letter).\nConsecutive kanji often expresses noun, and after it, a postpositional particle often appears. You can see some に,も,を,の appears after consecutive kanji. These are postpositional particles. \nI'm Japanese and I cannot read kuzushiji, but I can estimate the postpositional particle will come after difficult kanjis. I cannot tell which postpositional particle come up, but with image of the letter, I can predict it with some confidence.</p>\n\n<p>So I think the language model works with the help of the image recognition.\nI'm looking forward to seeing your great kernel :)</p>",
      "rawMarkdown": "&gt; do you think a language model should help here?\n\nYes, I think so. but I've not built a model yet.\nLet me explain with the first sentence as follows. \n\n `[自序] [若]い[時]の[気強]に[己]やれと[思]ふた[細工]も[老武者]の\nかなしさは[息子]に[及]ず[浮世]を[裏]の[三畳]に[避]て[正風]の\n[俳諧]を[楽]しめども[根]が[職人]の[文盲]だけこそけれ`\n\nThe kanji are written in parentheses, the others are kana (common easy letter).\nConsecutive kanji often expresses noun, and after it, a postpositional particle often appears. You can see some に,も,を,の appears after consecutive kanji. These are postpositional particles. \nI'm Japanese and I cannot read kuzushiji, but I can estimate the postpositional particle will come after difficult kanjis. I cannot tell which postpositional particle come up, but with image of the letter, I can predict it with some confidence.\n\nSo I think the language model works with the help of the image recognition.\nI'm looking forward to seeing your great kernel :)",
      "votes": 5,
      "replies": [
        {
          "id": 610206,
          "postDate": "2019-08-28T14:32:33.717Z",
          "content": "<p>Thank you for an explanation, I will try adding predictions of the language model to the image model and will share if this works.</p>",
          "rawMarkdown": "Thank you for an explanation, I will try adding predictions of the language model to the image model and will share if this works.",
          "votes": 2
        },
        {
          "id": 619420,
          "postDate": "2019-09-06T07:22:35.350Z",
          "content": "<p>Thanks @K_mat for your insight 😃 </p>",
          "rawMarkdown": "Thanks @K_mat for your insight 😃 "
        }
      ]
    },
    {
      "id": 620016,
      "postDate": "2019-09-06T21:54:27.917Z",
      "content": "<p>Potentially relevant post on the Google AI blog: <a href=\"https://ai.googleblog.com/2019/09/giving-lens-new-reading-capabilities-in.html\">https://ai.googleblog.com/2019/09/giving-lens-new-reading-capabilities-in.html</a></p>",
      "rawMarkdown": "Potentially relevant post on the Google AI blog: https://ai.googleblog.com/2019/09/giving-lens-new-reading-capabilities-in.html",
      "votes": 1
    },
    {
      "id": 610476,
      "postDate": "2019-08-28T20:33:12.610Z",
      "content": "<p>Personally I am struggling more with selecting the correct characters than classifying them correctly. I haven't tried language modelling, but I would speculate that the main benefit would be in correcting the characters with few/no examples in the training data, if you can find some good external data.</p>",
      "rawMarkdown": "Personally I am struggling more with selecting the correct characters than classifying them correctly. I haven't tried language modelling, but I would speculate that the main benefit would be in correcting the characters with few/no examples in the training data, if you can find some good external data.",
      "votes": 1,
      "replies": [
        {
          "id": 612373,
          "postDate": "2019-08-29T18:43:04.603Z",
          "content": "<blockquote>\n  <p>Personally I am struggling more with selecting the correct characters than classifying them correctly.</p>\n</blockquote>\n\n<p>This is interesting, do you mean you struggle more with detecting where are the characters on the page vs. classifying which of the characters it is?\nFor me it seems the main challenge is classification, while detection is quite accurate (although it could be improved as well).</p>",
          "rawMarkdown": "&gt; Personally I am struggling more with selecting the correct characters than classifying them correctly.\n\nThis is interesting, do you mean you struggle more with detecting where are the characters on the page vs. classifying which of the characters it is?\nFor me it seems the main challenge is classification, while detection is quite accurate (although it could be improved as well).",
          "votes": 2
        },
        {
          "id": 613836,
          "postDate": "2019-08-30T21:04:23.320Z",
          "content": "<p>Classification accuracy on a randomly selected validation set (which is 500 of the train images) is over 97%. I think my system for detecting where characters are contributes more to errors than the classifier, but that's just from looking at the results, I haven't calculated contributions to TP/FP/FN rates.</p>",
          "rawMarkdown": "Classification accuracy on a randomly selected validation set (which is 500 of the train images) is over 97%. I think my system for detecting where characters are contributes more to errors than the classifier, but that's just from looking at the results, I haven't calculated contributions to TP/FP/FN rates.",
          "votes": 3
        },
        {
          "id": 613875,
          "postDate": "2019-08-30T22:00:45.530Z",
          "content": "<p>I'm using the group split with book titles to evaluate the classification accuracy, as <a href=\"/lopuhin\">@lopuhin</a> mentioned. My current detection score (F1 socre of IOU&gt;0.5) is approx. 97% which is better than the classification accuracy of 92%.</p>",
          "rawMarkdown": "I'm using the group split with book titles to evaluate the classification accuracy, as @lopuhin mentioned. My current detection score (F1 socre of IOU&gt;0.5) is approx. 97% which is better than the classification accuracy of 92%.",
          "votes": 2
        },
        {
          "id": 613887,
          "postDate": "2019-08-30T22:44:05.463Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 619689,
          "postDate": "2019-09-06T13:18:40.317Z",
          "content": "<p>I might have misunderstood what you meant by \"group split with book titles\". Do you mean that each book should only appear in either the train or validation set?</p>",
          "rawMarkdown": "I might have misunderstood what you meant by \"group split with book titles\". Do you mean that each book should only appear in either the train or validation set?"
        },
        {
          "id": 619944,
          "postDate": "2019-09-06T19:44:26.823Z",
          "content": "<blockquote>\n  <p>Do you mean that each book should only appear in either the train or validation set?</p>\n</blockquote>\n\n<p>Yes, that's correct. We don't know for sure how the split between train and validation was done (right?), so it's a good point that we should not blindly assume this. But splitting by book seems quite natural, also if the split was done not by book, that would mean that there are some pages from books currently published on <a href=\"http://codh.rois.ac.jp/\">http://codh.rois.ac.jp/</a> which are missing and which are part of the test set.</p>",
          "rawMarkdown": "&gt; Do you mean that each book should only appear in either the train or validation set?\n\nYes, that's correct. We don't know for sure how the split between train and validation was done (right?), so it's a good point that we should not blindly assume this. But splitting by book seems quite natural, also if the split was done not by book, that would mean that there are some pages from books currently published on http://codh.rois.ac.jp/ which are missing and which are part of the test set.",
          "votes": 1
        },
        {
          "id": 619974,
          "postDate": "2019-09-06T20:33:03.747Z",
          "content": "<p>Thanks for information. To correct what I said above, my classification accuracy with a group split is 94% on the validation set (including a \"background\" class).</p>",
          "rawMarkdown": "Thanks for information. To correct what I said above, my classification accuracy with a group split is 94% on the validation set (including a \"background\" class)."
        },
        {
          "id": 620185,
          "postDate": "2019-09-07T06:03:51.640Z",
          "content": "<p>Thanks Ollie, that's really high! For me accuracy (also including background class, but it's only a few percent) is 0.87 -- 0.90 depending on the fold, so plenty of room to grow :)</p>",
          "rawMarkdown": "Thanks Ollie, that's really high! For me accuracy (also including background class, but it's only a few percent) is 0.87 -- 0.90 depending on the fold, so plenty of room to grow :)"
        }
      ]
    },
    {
      "id": 609880,
      "postDate": "2019-08-28T07:41:07Z",
      "content": "<p>How did you build the simple LSTM language model?  </p>",
      "rawMarkdown": "How did you build the simple LSTM language model?  ",
      "votes": 1,
      "replies": [
        {
          "id": 610006,
          "postDate": "2019-08-28T10:10:35.680Z",
          "content": "<p>The model was trained on text extracted from training data, with once symbol as one token. For that we need to extract symbol sequences from each page, here is an example of how this can be done: <a href=\"https://www.kaggle.com/rio114/order-charactors-by-clustering-columns\">https://www.kaggle.com/rio114/order-charactors-by-clustering-columns</a> (I used a different approach, but end result is similar).</p>\n\n<p>This is the model definition:\n```\nclass Model(nn.Module):\n    def <strong>init</strong>(\n            self, *,\n            n_classes: int,\n            embedding_dim: int = 256,\n            hidden_dim: int = 256,\n            ):\n        super().<strong>init</strong>()\n        self.embedding = nn.Embedding(n_classes, embedding_dim)\n        self.lstm = nn.LSTM(embedding_dim, hidden_dim, batch_first=True)\n        self.dropout = nn.Dropout(0.5)\n        self.fc_out = nn.Linear(hidden_dim, n_classes)</p>\n\n<pre><code>def forward(self, x):\n    x = self.embedding(x)\n    x, _ = self.lstm(x)\n    x = self.dropout(x)\n    x = self.fc_out(x)\n    return x.transpose(1, 2)\n</code></pre>\n\n<p>```</p>",
          "rawMarkdown": "The model was trained on text extracted from training data, with once symbol as one token. For that we need to extract symbol sequences from each page, here is an example of how this can be done: https://www.kaggle.com/rio114/order-charactors-by-clustering-columns (I used a different approach, but end result is similar).\n\nThis is the model definition:\n```\nclass Model(nn.Module):\n    def __init__(\n            self, *,\n            n_classes: int,\n            embedding_dim: int = 256,\n            hidden_dim: int = 256,\n            ):\n        super().__init__()\n        self.embedding = nn.Embedding(n_classes, embedding_dim)\n        self.lstm = nn.LSTM(embedding_dim, hidden_dim, batch_first=True)\n        self.dropout = nn.Dropout(0.5)\n        self.fc_out = nn.Linear(hidden_dim, n_classes)\n\n    def forward(self, x):\n        x = self.embedding(x)\n        x, _ = self.lstm(x)\n        x = self.dropout(x)\n        x = self.fc_out(x)\n        return x.transpose(1, 2)\n```",
          "votes": 5
        }
      ]
    },
    {
      "id": 642606,
      "postDate": "2019-10-06T11:24:22.160Z",
      "content": "<p>Can be a useful website too\n<a href=\"http://opac.lib.hiroshima-u.ac.jp/portal/dc/kyodo/naraehon/muromachi_top.html\">http://opac.lib.hiroshima-u.ac.jp/portal/dc/kyodo/naraehon/muromachi_top.html</a></p>",
      "rawMarkdown": "Can be a useful website too\nhttp://opac.lib.hiroshima-u.ac.jp/portal/dc/kyodo/naraehon/muromachi_top.html",
      "votes": 2
    },
    {
      "id": 610010,
      "postDate": "2019-08-28T10:12:16.633Z",
      "content": "<p>Maybe it would help to use some extra text data to train the language model, but I'm not sure where to get it so that it is similar enough to competition texts.</p>",
      "rawMarkdown": "Maybe it would help to use some extra text data to train the language model, but I'm not sure where to get it so that it is similar enough to competition texts.",
      "votes": 2,
      "replies": [
        {
          "id": 612879,
          "postDate": "2019-08-30T04:00:56.613Z",
          "rawMarkdown": "",
          "votes": 3,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 647792,
      "postDate": "2019-10-13T09:35:40.213Z",
      "content": "<p>Can't we just detect the characters as objects, where each character represents an object class?\nWhy do we even need to \"understand\" the writings, using a language model or whatever?</p>",
      "rawMarkdown": "Can't we just detect the characters as objects, where each character represents an object class?\nWhy do we even need to \"understand\" the writings, using a language model or whatever?",
      "replies": [
        {
          "id": 648022,
          "postDate": "2019-10-13T16:24:37.260Z",
          "content": "<p>Characters are often ambiguous without surrounding context. Example:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F885887%2Fbc0db3eca709f5099971b925853a5e6c%2Fte.png?generation=1570983705200695&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F885887%2Fad4bc17a7d2df4bd3c6ab4f4b84da61b%2Fku.png?generation=1570983717668700&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F885887%2Fc5a3dc295561c1b5a912698dd476ef6c%2Ftall_ku.png?generation=1570983731403289&amp;alt=media\" alt=\"\">\nMany of these are resolvable with the surrounding visual context, but some are not.</p>",
          "rawMarkdown": "Characters are often ambiguous without surrounding context. Example:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F885887%2Fbc0db3eca709f5099971b925853a5e6c%2Fte.png?generation=1570983705200695&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F885887%2Fad4bc17a7d2df4bd3c6ab4f4b84da61b%2Fku.png?generation=1570983717668700&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F885887%2Fc5a3dc295561c1b5a912698dd476ef6c%2Ftall_ku.png?generation=1570983731403289&amp;alt=media)\nMany of these are resolvable with the surrounding visual context, but some are not.",
          "votes": 2
        },
        {
          "id": 648075,
          "postDate": "2019-10-13T17:47:58.957Z",
          "content": "<p>As snowflower says, detection is certainly possible without context just by looking at the characters individually, but language context can help resolve ambiguities between similar (or poorly written) characters.</p>\n\n<p>For example, if I wrote the letter \"o\" on a page, it might be ambiguous whether the character is actually 'o' or '0'. But if you knew the context was \"leaderb_ard\", it becomes obvious the missing character is 'o' and not '0' (in fact, this is obvious even if the character 'o' is completely unreadable!)</p>",
          "rawMarkdown": "As snowflower says, detection is certainly possible without context just by looking at the characters individually, but language context can help resolve ambiguities between similar (or poorly written) characters.\n\nFor example, if I wrote the letter \"o\" on a page, it might be ambiguous whether the character is actually 'o' or '0'. But if you knew the context was \"leaderb_ard\", it becomes obvious the missing character is 'o' and not '0' (in fact, this is obvious even if the character 'o' is completely unreadable!)",
          "votes": 4
        }
      ]
    },
    {
      "id": 642602,
      "postDate": "2019-10-06T11:21:56.020Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 610096,
      "author_name": "K_mat",
      "author_url": "",
      "post_date": "2019-08-28T12:15:00.047000",
      "content": "<blockquote>\n  <p>do you think a language model should help here?</p>\n</blockquote>\n\n<p>Yes, I think so. but I've not built a model yet.\nLet me explain with the first sentence as follows. </p>\n\n<p><code>[自序] [若]い[時]の[気強]に[己]やれと[思]ふた[細工]も[老武者]の\nかなしさは[息子]に[及]ず[浮世]を[裏]の[三畳]に[避]て[正風]の\n[俳諧]を[楽]しめども[根]が[職人]の[文盲]だけこそけれ</code></p>\n\n<p>The kanji are written in parentheses, the others are kana (common easy letter).\nConsecutive kanji often expresses noun, and after it, a postpositional particle often appears. You can see some に,も,を,の appears after consecutive kanji. These are postpositional particles. \nI'm Japanese and I cannot read kuzushiji, but I can estimate the postpositional particle will come after difficult kanjis. I cannot tell which postpositional particle come up, but with image of the letter, I can predict it with some confidence.</p>\n\n<p>So I think the language model works with the help of the image recognition.\nI'm looking forward to seeing your great kernel :)</p>",
      "votes": 5,
      "replies": [
        {
          "id": 610206,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2019-08-28T14:32:33.717000",
          "content": "<p>Thank you for an explanation, I will try adding predictions of the language model to the image model and will share if this works.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 619420,
          "author_name": "Marrick Lip",
          "author_url": "",
          "post_date": "2019-09-06T07:22:35.350000",
          "content": "<p>Thanks @K_mat for your insight 😃 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 620016,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2019-09-06T21:54:27.917000",
      "content": "<p>Potentially relevant post on the Google AI blog: <a href=\"https://ai.googleblog.com/2019/09/giving-lens-new-reading-capabilities-in.html\">https://ai.googleblog.com/2019/09/giving-lens-new-reading-capabilities-in.html</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 610476,
      "author_name": "Ollie Perrée",
      "author_url": "",
      "post_date": "2019-08-28T20:33:12.610000",
      "content": "<p>Personally I am struggling more with selecting the correct characters than classifying them correctly. I haven't tried language modelling, but I would speculate that the main benefit would be in correcting the characters with few/no examples in the training data, if you can find some good external data.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 612373,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2019-08-29T18:43:04.603000",
          "content": "<blockquote>\n  <p>Personally I am struggling more with selecting the correct characters than classifying them correctly.</p>\n</blockquote>\n\n<p>This is interesting, do you mean you struggle more with detecting where are the characters on the page vs. classifying which of the characters it is?\nFor me it seems the main challenge is classification, while detection is quite accurate (although it could be improved as well).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 613836,
          "author_name": "Ollie Perrée",
          "author_url": "",
          "post_date": "2019-08-30T21:04:23.320000",
          "content": "<p>Classification accuracy on a randomly selected validation set (which is 500 of the train images) is over 97%. I think my system for detecting where characters are contributes more to errors than the classifier, but that's just from looking at the results, I haven't calculated contributions to TP/FP/FN rates.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 613875,
          "author_name": "K_mat",
          "author_url": "",
          "post_date": "2019-08-30T22:00:45.530000",
          "content": "<p>I'm using the group split with book titles to evaluate the classification accuracy, as <a href=\"/lopuhin\">@lopuhin</a> mentioned. My current detection score (F1 socre of IOU&gt;0.5) is approx. 97% which is better than the classification accuracy of 92%.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 613887,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-08-30T22:44:05.463000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 619689,
          "author_name": "Ollie Perrée",
          "author_url": "",
          "post_date": "2019-09-06T13:18:40.317000",
          "content": "<p>I might have misunderstood what you meant by \"group split with book titles\". Do you mean that each book should only appear in either the train or validation set?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 619944,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2019-09-06T19:44:26.823000",
          "content": "<blockquote>\n  <p>Do you mean that each book should only appear in either the train or validation set?</p>\n</blockquote>\n\n<p>Yes, that's correct. We don't know for sure how the split between train and validation was done (right?), so it's a good point that we should not blindly assume this. But splitting by book seems quite natural, also if the split was done not by book, that would mean that there are some pages from books currently published on <a href=\"http://codh.rois.ac.jp/\">http://codh.rois.ac.jp/</a> which are missing and which are part of the test set.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 619974,
          "author_name": "Ollie Perrée",
          "author_url": "",
          "post_date": "2019-09-06T20:33:03.747000",
          "content": "<p>Thanks for information. To correct what I said above, my classification accuracy with a group split is 94% on the validation set (including a \"background\" class).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 620185,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2019-09-07T06:03:51.640000",
          "content": "<p>Thanks Ollie, that's really high! For me accuracy (also including background class, but it's only a few percent) is 0.87 -- 0.90 depending on the fold, so plenty of room to grow :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 609880,
      "author_name": "TheNuttyNetter",
      "author_url": "",
      "post_date": "2019-08-28T07:41:07",
      "content": "<p>How did you build the simple LSTM language model?  </p>",
      "votes": 1,
      "replies": [
        {
          "id": 610006,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2019-08-28T10:10:35.680000",
          "content": "<p>The model was trained on text extracted from training data, with once symbol as one token. For that we need to extract symbol sequences from each page, here is an example of how this can be done: <a href=\"https://www.kaggle.com/rio114/order-charactors-by-clustering-columns\">https://www.kaggle.com/rio114/order-charactors-by-clustering-columns</a> (I used a different approach, but end result is similar).</p>\n\n<p>This is the model definition:\n```\nclass Model(nn.Module):\n    def <strong>init</strong>(\n            self, *,\n            n_classes: int,\n            embedding_dim: int = 256,\n            hidden_dim: int = 256,\n            ):\n        super().<strong>init</strong>()\n        self.embedding = nn.Embedding(n_classes, embedding_dim)\n        self.lstm = nn.LSTM(embedding_dim, hidden_dim, batch_first=True)\n        self.dropout = nn.Dropout(0.5)\n        self.fc_out = nn.Linear(hidden_dim, n_classes)</p>\n\n<pre><code>def forward(self, x):\n    x = self.embedding(x)\n    x, _ = self.lstm(x)\n    x = self.dropout(x)\n    x = self.fc_out(x)\n    return x.transpose(1, 2)\n</code></pre>\n\n<p>```</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 642606,
      "author_name": "Nadezda Bertseva",
      "author_url": "",
      "post_date": "2019-10-06T11:24:22.160000",
      "content": "<p>Can be a useful website too\n<a href=\"http://opac.lib.hiroshima-u.ac.jp/portal/dc/kyodo/naraehon/muromachi_top.html\">http://opac.lib.hiroshima-u.ac.jp/portal/dc/kyodo/naraehon/muromachi_top.html</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 610010,
      "author_name": "Konstantin Lopukhin",
      "author_url": "",
      "post_date": "2019-08-28T10:12:16.633000",
      "content": "<p>Maybe it would help to use some extra text data to train the language model, but I'm not sure where to get it so that it is similar enough to competition texts.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 612879,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-08-30T04:00:56.613000",
          "content": "",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 647792,
      "author_name": "Atom",
      "author_url": "",
      "post_date": "2019-10-13T09:35:40.213000",
      "content": "<p>Can't we just detect the characters as objects, where each character represents an object class?\nWhy do we even need to \"understand\" the writings, using a language model or whatever?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 648022,
          "author_name": ">///<",
          "author_url": "",
          "post_date": "2019-10-13T16:24:37.260000",
          "content": "<p>Characters are often ambiguous without surrounding context. Example:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F885887%2Fbc0db3eca709f5099971b925853a5e6c%2Fte.png?generation=1570983705200695&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F885887%2Fad4bc17a7d2df4bd3c6ab4f4b84da61b%2Fku.png?generation=1570983717668700&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F885887%2Fc5a3dc295561c1b5a912698dd476ef6c%2Ftall_ku.png?generation=1570983731403289&amp;alt=media\" alt=\"\">\nMany of these are resolvable with the surrounding visual context, but some are not.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 648075,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "2019-10-13T17:47:58.957000",
          "content": "<p>As snowflower says, detection is certainly possible without context just by looking at the characters individually, but language context can help resolve ambiguities between similar (or poorly written) characters.</p>\n\n<p>For example, if I wrote the letter \"o\" on a page, it might be ambiguous whether the character is actually 'o' or '0'. But if you knew the context was \"leaderb_ard\", it becomes obvious the missing character is 'o' and not '0' (in fact, this is obvious even if the character 'o' is completely unreadable!)</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 642602,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-10-06T11:21:56.020000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "609876": "I wonder if anyone was able to successfully use a language model? It should be helpful because as I understand, in many cases a character would be easier to recognize knowing it's context, and this is a technique commonly used with OCR, see for example https://medium.com/apache-mxnet/handwriting-ocr-handwriting-recognition-and-language-modeling-with-mxnet-gluon-4c7165788c67\nI tried building a simple LSTM language model, but the best validation loss (cross-entropy) I could get was 4.8 (with a per-book validation split), which sounds extremely bad compared to 0.6402 cross-entropy loss I get with the same split with the image model (which gets my current best LB score of 0.863, it's a single fold single model).\n\nI'm interested in your thoughts on the following:\n- do you think a language model should help here?\n- did you try building a language model and if yes, what kind of log loss / perplexity did you get?\n- did you try incorporating the language model into your solution, and if yes, did it improve the score?",
    "610096": "&gt; do you think a language model should help here?\n\nYes, I think so. but I've not built a model yet.\nLet me explain with the first sentence as follows. \n\n `[自序] [若]い[時]の[気強]に[己]やれと[思]ふた[細工]も[老武者]の\nかなしさは[息子]に[及]ず[浮世]を[裏]の[三畳]に[避]て[正風]の\n[俳諧]を[楽]しめども[根]が[職人]の[文盲]だけこそけれ`\n\nThe kanji are written in parentheses, the others are kana (common easy letter).\nConsecutive kanji often expresses noun, and after it, a postpositional particle often appears. You can see some に,も,を,の appears after consecutive kanji. These are postpositional particles. \nI'm Japanese and I cannot read kuzushiji, but I can estimate the postpositional particle will come after difficult kanjis. I cannot tell which postpositional particle come up, but with image of the letter, I can predict it with some confidence.\n\nSo I think the language model works with the help of the image recognition.\nI'm looking forward to seeing your great kernel :)",
    "620016": "Potentially relevant post on the Google AI blog: https://ai.googleblog.com/2019/09/giving-lens-new-reading-capabilities-in.html",
    "610476": "Personally I am struggling more with selecting the correct characters than classifying them correctly. I haven't tried language modelling, but I would speculate that the main benefit would be in correcting the characters with few/no examples in the training data, if you can find some good external data.",
    "609880": "How did you build the simple LSTM language model?  ",
    "642606": "Can be a useful website too\nhttp://opac.lib.hiroshima-u.ac.jp/portal/dc/kyodo/naraehon/muromachi_top.html",
    "610010": "Maybe it would help to use some extra text data to train the language model, but I'm not sure where to get it so that it is similar enough to competition texts.",
    "647792": "Can't we just detect the characters as objects, where each character represents an object class?\nWhy do we even need to \"understand\" the writings, using a language model or whatever?",
    "642602": ""
  }
}