{
  "id": 114764,
  "title": "4th place solution",
  "url": "/competitions/kuzushiji-recognition/discussion/114764",
  "author_name": "kolman",
  "post_date": "2019-10-29T04:06:55.704000",
  "votes": 21,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Thanks to all for organizers this very interesting competition.\nCongratulations to all who finished the competition and to the winners.</p>\n\n<p>Below is an overview of my solution.\ncode: <a href=\"https://github.com/linhuifj/kaggle-kuzushiji-recognition\">https://github.com/linhuifj/kaggle-kuzushiji-recognition</a>\nmethod:\ncharacter detection -&gt; get lines -&gt; line recognition -&gt; postprocessing</p>\n\n<p><strong>Detection</strong>\nA detector is trained to find all words in the input image. Here <a href=\"http://openaccess.thecvf.com/content_CVPR_2019/html/Chen_Hybrid_Task_Cascade_for_Instance_Segmentation_CVPR_2019_paper.html\">HTC(Hybrid Task Cascade)</a> is adopted, which is the SOTA two-stage object detection method;</p>\n\n<p>Words are grouped into different lines, using the text line constructing algorithm used by CTPN;</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F237552%2F4d33c3040ac8c4f8fe7835afdd3b0e2a%2Fclipboard.png?generation=1572320037825623&amp;alt=media\" alt=\"\"></p>\n\n<p>Lines are cropped from the input image for line-wise recognition.\nWhen cropping the lines out, each pixel coordinate of \"line top\" line(the red line in the image below) is recorded for mapping the predicted positions of each word in the cropped line back to the whole page image.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F237552%2F1b7b1d608c2a649e298d30998eba2ad9%2Fclipboard1.png?generation=1572320223576382&amp;alt=media\" alt=\"\"></p>\n\n<p>The cropped line images are all non-slant as shown below.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F237552%2F1ac013f6a815931f68b3bb546ea01c2e%2F100241706_00004_2_l_3.jpg?generation=1572320293387444&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>Recognition</strong>\nThe lines extracted from the detection stage are resized (and padded) to dimension 32x800, and fed into a line recognition model.\nWe take the CRNN model for line recognition.  For each line, the model has 200 outputs. \nWe trained a six-gram language model by KENLM, and use beam-search to decode the CTC output of the model. \nThe position of each character is calculated by taking account of the coordinate of each line in the original image and the output of CTC.</p>\n\n<p><strong>Post-processing</strong>\nWe binarize each image to get the stroke of the words. For each center point calculated by the recognition stage, we find the nearest stroke and set the center point to the coordinate of the found position. This is called \"center point adjustment\".</p>\n\n<p>Because our lines may overlap and generate duplicate results, we use another script to remove the duplicate outputs.</p>\n\n<p><strong>Interesting Findings</strong>\n1. Position Accuracy\nUse multitask training by adding attention loss will increase the position accuracy of the CTC output. Refer to \"JOINT CTC-ATTENTION BASED END TO-END SPEECH RECOGNITION USING MULTI-TASK LEARNING\" for model architecture. We found that the output of CTC has a smaller edit distance while the output of attention has higher line accuracy, but smaller edit distance is better for this task.\nWe cannot use more than 2 layers of LSTM or the position of the result will drift and become inaccurate.</p>\n\n<ol>\n<li><p>Data augmentation\nData augmentation is very important for line recognition model. We use \nCLAHE, Solarize, RandomBrightness, RandomContrast, random scale, random distort for data augmentation. \nThe augmentated lines are shown as follows:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F237552%2F4e574bbe1ef4b87e335569def9c6383a%2Faug0.jpg?generation=1572325026371505&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F237552%2Fc9698008de3141dd211e902e3dba3a9e%2Faug1.jpg?generation=1572325076407727&amp;alt=media\" alt=\"\"></p></li>\n<li><p>Regularization\nDropout is added between the last layer of cnn and the first layer of lstm.\nCutout is also used. The model overfits and oscillates in performance so early stopping is also very important.</p></li>\n<li><p>Language model \nLanguage Model can make a slight improvement in the final score. </p></li>\n<li><p>Model ensemble\nWe do not use model ensemble.</p></li>\n</ol>\n\n<p><strong>Logs of score</strong>\n| method | public leaderboard score |\n| --- | --- |\n|  baseline (no dropout, no data augmentation) |  0.846 |\n| vgg + ctc + lstm,  + data augmentation, no language model, no dropout| 0.877|\n| resnet + ctc | 0.885 |\n| + dropout | 0.905 | \n| + adjust center | 0.915|\n| + lstm + attention | 0.925 |\n| + language model | 0.928|\n| sgd finetune | 0.932|\n| + cutout | 0.935|\n| remove duplicate| 0.938|</p>\n\n<p><strong>Hardware</strong>\nFor detection network, we use 8xm40 for 5 days, and for recognition network, we use 8x2080ti to train for one week. </p>",
  "messages": [
    {
      "id": 660385,
      "postDate": "2019-10-29T04:06:55.703Z",
      "content": "<p>Thanks to all for organizers this very interesting competition.\nCongratulations to all who finished the competition and to the winners.</p>\n\n<p>Below is an overview of my solution.\ncode: <a href=\"https://github.com/linhuifj/kaggle-kuzushiji-recognition\">https://github.com/linhuifj/kaggle-kuzushiji-recognition</a>\nmethod:\ncharacter detection -&gt; get lines -&gt; line recognition -&gt; postprocessing</p>\n\n<p><strong>Detection</strong>\nA detector is trained to find all words in the input image. Here <a href=\"http://openaccess.thecvf.com/content_CVPR_2019/html/Chen_Hybrid_Task_Cascade_for_Instance_Segmentation_CVPR_2019_paper.html\">HTC(Hybrid Task Cascade)</a> is adopted, which is the SOTA two-stage object detection method;</p>\n\n<p>Words are grouped into different lines, using the text line constructing algorithm used by CTPN;</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F237552%2F4d33c3040ac8c4f8fe7835afdd3b0e2a%2Fclipboard.png?generation=1572320037825623&amp;alt=media\" alt=\"\"></p>\n\n<p>Lines are cropped from the input image for line-wise recognition.\nWhen cropping the lines out, each pixel coordinate of \"line top\" line(the red line in the image below) is recorded for mapping the predicted positions of each word in the cropped line back to the whole page image.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F237552%2F1b7b1d608c2a649e298d30998eba2ad9%2Fclipboard1.png?generation=1572320223576382&amp;alt=media\" alt=\"\"></p>\n\n<p>The cropped line images are all non-slant as shown below.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F237552%2F1ac013f6a815931f68b3bb546ea01c2e%2F100241706_00004_2_l_3.jpg?generation=1572320293387444&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>Recognition</strong>\nThe lines extracted from the detection stage are resized (and padded) to dimension 32x800, and fed into a line recognition model.\nWe take the CRNN model for line recognition.  For each line, the model has 200 outputs. \nWe trained a six-gram language model by KENLM, and use beam-search to decode the CTC output of the model. \nThe position of each character is calculated by taking account of the coordinate of each line in the original image and the output of CTC.</p>\n\n<p><strong>Post-processing</strong>\nWe binarize each image to get the stroke of the words. For each center point calculated by the recognition stage, we find the nearest stroke and set the center point to the coordinate of the found position. This is called \"center point adjustment\".</p>\n\n<p>Because our lines may overlap and generate duplicate results, we use another script to remove the duplicate outputs.</p>\n\n<p><strong>Interesting Findings</strong>\n1. Position Accuracy\nUse multitask training by adding attention loss will increase the position accuracy of the CTC output. Refer to \"JOINT CTC-ATTENTION BASED END TO-END SPEECH RECOGNITION USING MULTI-TASK LEARNING\" for model architecture. We found that the output of CTC has a smaller edit distance while the output of attention has higher line accuracy, but smaller edit distance is better for this task.\nWe cannot use more than 2 layers of LSTM or the position of the result will drift and become inaccurate.</p>\n\n<ol>\n<li><p>Data augmentation\nData augmentation is very important for line recognition model. We use \nCLAHE, Solarize, RandomBrightness, RandomContrast, random scale, random distort for data augmentation. \nThe augmentated lines are shown as follows:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F237552%2F4e574bbe1ef4b87e335569def9c6383a%2Faug0.jpg?generation=1572325026371505&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F237552%2Fc9698008de3141dd211e902e3dba3a9e%2Faug1.jpg?generation=1572325076407727&amp;alt=media\" alt=\"\"></p></li>\n<li><p>Regularization\nDropout is added between the last layer of cnn and the first layer of lstm.\nCutout is also used. The model overfits and oscillates in performance so early stopping is also very important.</p></li>\n<li><p>Language model \nLanguage Model can make a slight improvement in the final score. </p></li>\n<li><p>Model ensemble\nWe do not use model ensemble.</p></li>\n</ol>\n\n<p><strong>Logs of score</strong>\n| method | public leaderboard score |\n| --- | --- |\n|  baseline (no dropout, no data augmentation) |  0.846 |\n| vgg + ctc + lstm,  + data augmentation, no language model, no dropout| 0.877|\n| resnet + ctc | 0.885 |\n| + dropout | 0.905 | \n| + adjust center | 0.915|\n| + lstm + attention | 0.925 |\n| + language model | 0.928|\n| sgd finetune | 0.932|\n| + cutout | 0.935|\n| remove duplicate| 0.938|</p>\n\n<p><strong>Hardware</strong>\nFor detection network, we use 8xm40 for 5 days, and for recognition network, we use 8x2080ti to train for one week. </p>",
      "rawMarkdown": "Thanks to all for organizers this very interesting competition.\nCongratulations to all who finished the competition and to the winners.\n\nBelow is an overview of my solution.\ncode: https://github.com/linhuifj/kaggle-kuzushiji-recognition\nmethod:\ncharacter detection -&gt; get lines -&gt; line recognition -&gt; postprocessing\n\n**Detection**\nA detector is trained to find all words in the input image. Here [HTC(Hybrid Task Cascade)](http://openaccess.thecvf.com/content_CVPR_2019/html/Chen_Hybrid_Task_Cascade_for_Instance_Segmentation_CVPR_2019_paper.html) is adopted, which is the SOTA two-stage object detection method;\n\nWords are grouped into different lines, using the text line constructing algorithm used by CTPN;\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F237552%2F4d33c3040ac8c4f8fe7835afdd3b0e2a%2Fclipboard.png?generation=1572320037825623&amp;alt=media)\n\nLines are cropped from the input image for line-wise recognition.\nWhen cropping the lines out, each pixel coordinate of \"line top\" line(the red line in the image below) is recorded for mapping the predicted positions of each word in the cropped line back to the whole page image.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F237552%2F1b7b1d608c2a649e298d30998eba2ad9%2Fclipboard1.png?generation=1572320223576382&amp;alt=media)\n\nThe cropped line images are all non-slant as shown below.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F237552%2F1ac013f6a815931f68b3bb546ea01c2e%2F100241706_00004_2_l_3.jpg?generation=1572320293387444&amp;alt=media)\n\n**Recognition**\nThe lines extracted from the detection stage are resized (and padded) to dimension 32x800, and fed into a line recognition model.\nWe take the CRNN model for line recognition.  For each line, the model has 200 outputs. \nWe trained a six-gram language model by KENLM, and use beam-search to decode the CTC output of the model. \nThe position of each character is calculated by taking account of the coordinate of each line in the original image and the output of CTC.\n\n**Post-processing**\nWe binarize each image to get the stroke of the words. For each center point calculated by the recognition stage, we find the nearest stroke and set the center point to the coordinate of the found position. This is called \"center point adjustment\".\n\nBecause our lines may overlap and generate duplicate results, we use another script to remove the duplicate outputs.\n\n**Interesting Findings**\n1. Position Accuracy\nUse multitask training by adding attention loss will increase the position accuracy of the CTC output. Refer to \"JOINT CTC-ATTENTION BASED END TO-END SPEECH RECOGNITION USING MULTI-TASK LEARNING\" for model architecture. We found that the output of CTC has a smaller edit distance while the output of attention has higher line accuracy, but smaller edit distance is better for this task.\nWe cannot use more than 2 layers of LSTM or the position of the result will drift and become inaccurate.\n\n2. Data augmentation\nData augmentation is very important for line recognition model. We use \nCLAHE, Solarize, RandomBrightness, RandomContrast, random scale, random distort for data augmentation. \nThe augmentated lines are shown as follows:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F237552%2F4e574bbe1ef4b87e335569def9c6383a%2Faug0.jpg?generation=1572325026371505&amp;alt=media)\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F237552%2Fc9698008de3141dd211e902e3dba3a9e%2Faug1.jpg?generation=1572325076407727&amp;alt=media)\n\n3. Regularization\nDropout is added between the last layer of cnn and the first layer of lstm.\nCutout is also used. The model overfits and oscillates in performance so early stopping is also very important.\n\n4. Language model \nLanguage Model can make a slight improvement in the final score. \n\n5. Model ensemble\nWe do not use model ensemble.\n\n**Logs of score**\n| method | public leaderboard score |\n| --- | --- |\n|  baseline (no dropout, no data augmentation) |  0.846 |\n| vgg + ctc + lstm,  + data augmentation, no language model, no dropout| 0.877|\n| resnet + ctc | 0.885 |\n| + dropout | 0.905 | \n| + adjust center | 0.915|\n| + lstm + attention | 0.925 |\n| + language model | 0.928|\n| sgd finetune | 0.932|\n| + cutout | 0.935|\n| remove duplicate| 0.938|\n\n**Hardware**\nFor detection network, we use 8xm40 for 5 days, and for recognition network, we use 8x2080ti to train for one week. ",
      "votes": 21
    },
    {
      "id": 675393,
      "postDate": "2019-11-18T03:04:28.183Z",
      "content": "<p>Hello, </p>\n\n<p>Organizer here.  If you wouldn't mind, can you say how much time your method takes for inference per-page-image (even a ballpark or rough estimate is fine)?  </p>\n\n<p>Best, </p>\n\n<p>Alex.  </p>",
      "rawMarkdown": "Hello, \n\nOrganizer here.  If you wouldn't mind, can you say how much time your method takes for inference per-page-image (even a ballpark or rough estimate is fine)?  \n\nBest, \n\nAlex.  ",
      "votes": 1,
      "replies": [
        {
          "id": 704180,
          "postDate": "2019-12-27T06:23:47.010Z",
          "content": "<p>Less than 500ms. It could be much faster if we optimize the code and model.</p>",
          "rawMarkdown": "Less than 500ms. It could be much faster if we optimize the code and model."
        }
      ]
    },
    {
      "id": 966447,
      "postDate": "2020-08-11T12:24:41.583Z",
      "content": "<p>Hi. I'm curious to now that where did you get this 8x2080Ti GPU system from? Do you own this yourself?</p>",
      "rawMarkdown": "Hi. I'm curious to now that where did you get this 8x2080Ti GPU system from? Do you own this yourself?"
    },
    {
      "id": 814329,
      "postDate": "2020-04-20T15:29:00.080Z",
      "content": "<p>There are many better ocr detectors and recognitors , why you choose the base one CRNN?  </p>",
      "rawMarkdown": "There are many better ocr detectors and recognitors , why you choose the base one CRNN?  ",
      "replies": [
        {
          "id": 1852437,
          "postDate": "2022-07-12T04:34:02.157Z",
          "content": "<p>We choose CRNN because it is simple. Yes there are many better detectors and recognizers, so the result could be better.</p>",
          "rawMarkdown": "We choose CRNN because it is simple. Yes there are many better detectors and recognizers, so the result could be better."
        }
      ]
    },
    {
      "id": 754958,
      "postDate": "2020-02-24T09:17:31.773Z",
      "content": "<p>Great</p>",
      "rawMarkdown": "Great"
    },
    {
      "id": 665422,
      "postDate": "2019-11-05T01:49:39.793Z",
      "content": "<p>great!</p>",
      "rawMarkdown": "great!"
    },
    {
      "id": 660582,
      "postDate": "2019-10-29T11:05:04.327Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 675393,
      "author_name": "TheNuttyNetter",
      "author_url": "",
      "post_date": "2019-11-18T03:04:28.183000",
      "content": "<p>Hello, </p>\n\n<p>Organizer here.  If you wouldn't mind, can you say how much time your method takes for inference per-page-image (even a ballpark or rough estimate is fine)?  </p>\n\n<p>Best, </p>\n\n<p>Alex.  </p>",
      "votes": 1,
      "replies": [
        {
          "id": 704180,
          "author_name": "kolman",
          "author_url": "",
          "post_date": "2019-12-27T06:23:47.010000",
          "content": "<p>Less than 500ms. It could be much faster if we optimize the code and model.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 966447,
      "author_name": "Gaurav",
      "author_url": "",
      "post_date": "2020-08-11T12:24:41.583000",
      "content": "<p>Hi. I'm curious to now that where did you get this 8x2080Ti GPU system from? Do you own this yourself?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 814329,
      "author_name": "Bryce1010",
      "author_url": "",
      "post_date": "2020-04-20T15:29:00.080000",
      "content": "<p>There are many better ocr detectors and recognitors , why you choose the base one CRNN?  </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1852437,
          "author_name": "kolman",
          "author_url": "",
          "post_date": "2022-07-12T04:34:02.157000",
          "content": "<p>We choose CRNN because it is simple. Yes there are many better detectors and recognizers, so the result could be better.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 754958,
      "author_name": "Rishabh Shukla",
      "author_url": "",
      "post_date": "2020-02-24T09:17:31.773000",
      "content": "<p>Great</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 665422,
      "author_name": "bingfeng",
      "author_url": "",
      "post_date": "2019-11-05T01:49:39.793000",
      "content": "<p>great!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 660582,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-10-29T11:05:04.327000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "660385": "Thanks to all for organizers this very interesting competition.\nCongratulations to all who finished the competition and to the winners.\n\nBelow is an overview of my solution.\ncode: https://github.com/linhuifj/kaggle-kuzushiji-recognition\nmethod:\ncharacter detection -&gt; get lines -&gt; line recognition -&gt; postprocessing\n\n**Detection**\nA detector is trained to find all words in the input image. Here [HTC(Hybrid Task Cascade)](http://openaccess.thecvf.com/content_CVPR_2019/html/Chen_Hybrid_Task_Cascade_for_Instance_Segmentation_CVPR_2019_paper.html) is adopted, which is the SOTA two-stage object detection method;\n\nWords are grouped into different lines, using the text line constructing algorithm used by CTPN;\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F237552%2F4d33c3040ac8c4f8fe7835afdd3b0e2a%2Fclipboard.png?generation=1572320037825623&amp;alt=media)\n\nLines are cropped from the input image for line-wise recognition.\nWhen cropping the lines out, each pixel coordinate of \"line top\" line(the red line in the image below) is recorded for mapping the predicted positions of each word in the cropped line back to the whole page image.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F237552%2F1b7b1d608c2a649e298d30998eba2ad9%2Fclipboard1.png?generation=1572320223576382&amp;alt=media)\n\nThe cropped line images are all non-slant as shown below.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F237552%2F1ac013f6a815931f68b3bb546ea01c2e%2F100241706_00004_2_l_3.jpg?generation=1572320293387444&amp;alt=media)\n\n**Recognition**\nThe lines extracted from the detection stage are resized (and padded) to dimension 32x800, and fed into a line recognition model.\nWe take the CRNN model for line recognition.  For each line, the model has 200 outputs. \nWe trained a six-gram language model by KENLM, and use beam-search to decode the CTC output of the model. \nThe position of each character is calculated by taking account of the coordinate of each line in the original image and the output of CTC.\n\n**Post-processing**\nWe binarize each image to get the stroke of the words. For each center point calculated by the recognition stage, we find the nearest stroke and set the center point to the coordinate of the found position. This is called \"center point adjustment\".\n\nBecause our lines may overlap and generate duplicate results, we use another script to remove the duplicate outputs.\n\n**Interesting Findings**\n1. Position Accuracy\nUse multitask training by adding attention loss will increase the position accuracy of the CTC output. Refer to \"JOINT CTC-ATTENTION BASED END TO-END SPEECH RECOGNITION USING MULTI-TASK LEARNING\" for model architecture. We found that the output of CTC has a smaller edit distance while the output of attention has higher line accuracy, but smaller edit distance is better for this task.\nWe cannot use more than 2 layers of LSTM or the position of the result will drift and become inaccurate.\n\n2. Data augmentation\nData augmentation is very important for line recognition model. We use \nCLAHE, Solarize, RandomBrightness, RandomContrast, random scale, random distort for data augmentation. \nThe augmentated lines are shown as follows:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F237552%2F4e574bbe1ef4b87e335569def9c6383a%2Faug0.jpg?generation=1572325026371505&amp;alt=media)\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F237552%2Fc9698008de3141dd211e902e3dba3a9e%2Faug1.jpg?generation=1572325076407727&amp;alt=media)\n\n3. Regularization\nDropout is added between the last layer of cnn and the first layer of lstm.\nCutout is also used. The model overfits and oscillates in performance so early stopping is also very important.\n\n4. Language model \nLanguage Model can make a slight improvement in the final score. \n\n5. Model ensemble\nWe do not use model ensemble.\n\n**Logs of score**\n| method | public leaderboard score |\n| --- | --- |\n|  baseline (no dropout, no data augmentation) |  0.846 |\n| vgg + ctc + lstm,  + data augmentation, no language model, no dropout| 0.877|\n| resnet + ctc | 0.885 |\n| + dropout | 0.905 | \n| + adjust center | 0.915|\n| + lstm + attention | 0.925 |\n| + language model | 0.928|\n| sgd finetune | 0.932|\n| + cutout | 0.935|\n| remove duplicate| 0.938|\n\n**Hardware**\nFor detection network, we use 8xm40 for 5 days, and for recognition network, we use 8x2080ti to train for one week. ",
    "675393": "Hello, \n\nOrganizer here.  If you wouldn't mind, can you say how much time your method takes for inference per-page-image (even a ballpark or rough estimate is fine)?  \n\nBest, \n\nAlex.  ",
    "966447": "Hi. I'm curious to now that where did you get this 8x2080Ti GPU system from? Do you own this yourself?",
    "814329": "There are many better ocr detectors and recognitors , why you choose the base one CRNN?  ",
    "754958": "Great",
    "665422": "great!",
    "660582": ""
  }
}