{
  "id": 231400,
  "title": "Making inferencing faster",
  "url": "/competitions/bms-molecular-translation/discussion/231400",
  "author_name": "tugstugi",
  "post_date": "2021-04-08T11:49:42.296000",
  "votes": 26,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Inferencing on 1.5 million images is really slow. So i made some experiments with resnet34 to accelerate the inferencing. Here are results on the validation set (20% of the train set):</p>\n<pre><code>GPU: 1x 2080ti\nencoder: resnet34\ndecoder: lstm with attention\nbatch size: 256\nworker: 8\nimage size: 224x224\n\n20m14s: inferencing like Y.Nakamas public notebook\n16m50s: break if batch EOS\n12m39s: break if batch EOS + sort by InChi length\n12m42s: break if batch EOS + sort by InChi length + torch.jit.script(encoder)\n</code></pre>\n<p>So with sorting by InChi length, you can make your inferencing 35% faster if you have already a previous submission. If not, you can use batch EOS break which is still much faster than the naive inferencing.</p>\n<p>To use the batch EOS break, change the Y.Nakamas code this way:</p>\n<pre><code>for t in range(decode_lengths):\n    ...\n    if np.argmax(preds.detach().cpu().numpy()) == tokenizer.stoi[\"&lt;eos&gt;\"]:\n        break\n</code></pre>\n<p>to</p>\n<pre><code>end_condition = torch.zeros(batch_size, dtype=torch.long).to(encoder_out.device)\nfor t in range(decode_lengths):\n   ...\n   end_condition |= (torch.argmax(preds, -1) == tokenizer.stoi[\"&lt;eos&gt;\"])\n       if end_condition.sum() == batch_size:\n           break\n</code></pre>\n<p>Using torch.jit on the resnet34 encoder seems to have no effect.</p>",
  "messages": [
    {
      "id": 1267215,
      "postDate": "2021-04-08T11:49:42.297Z",
      "content": "<p>Inferencing on 1.5 million images is really slow. So i made some experiments with resnet34 to accelerate the inferencing. Here are results on the validation set (20% of the train set):</p>\n<pre><code>GPU: 1x 2080ti\nencoder: resnet34\ndecoder: lstm with attention\nbatch size: 256\nworker: 8\nimage size: 224x224\n\n20m14s: inferencing like Y.Nakamas public notebook\n16m50s: break if batch EOS\n12m39s: break if batch EOS + sort by InChi length\n12m42s: break if batch EOS + sort by InChi length + torch.jit.script(encoder)\n</code></pre>\n<p>So with sorting by InChi length, you can make your inferencing 35% faster if you have already a previous submission. If not, you can use batch EOS break which is still much faster than the naive inferencing.</p>\n<p>To use the batch EOS break, change the Y.Nakamas code this way:</p>\n<pre><code>for t in range(decode_lengths):\n    ...\n    if np.argmax(preds.detach().cpu().numpy()) == tokenizer.stoi[\"&lt;eos&gt;\"]:\n        break\n</code></pre>\n<p>to</p>\n<pre><code>end_condition = torch.zeros(batch_size, dtype=torch.long).to(encoder_out.device)\nfor t in range(decode_lengths):\n   ...\n   end_condition |= (torch.argmax(preds, -1) == tokenizer.stoi[\"&lt;eos&gt;\"])\n       if end_condition.sum() == batch_size:\n           break\n</code></pre>\n<p>Using torch.jit on the resnet34 encoder seems to have no effect.</p>",
      "rawMarkdown": "Inferencing on 1.5 million images is really slow. So i made some experiments with resnet34 to accelerate the inferencing. Here are results on the validation set (20% of the train set):\n```\nGPU: 1x 2080ti\nencoder: resnet34\ndecoder: lstm with attention\nbatch size: 256\nworker: 8\nimage size: 224x224\n\n20m14s: inferencing like Y.Nakamas public notebook\n16m50s: break if batch EOS\n12m39s: break if batch EOS + sort by InChi length\n12m42s: break if batch EOS + sort by InChi length + torch.jit.script(encoder)\n```\n\nSo with sorting by InChi length, you can make your inferencing 35% faster if you have already a previous submission. If not, you can use batch EOS break which is still much faster than the naive inferencing.\n\nTo use the batch EOS break, change the Y.Nakamas code this way:\n```\nfor t in range(decode_lengths):\n    ...\n    if np.argmax(preds.detach().cpu().numpy()) == tokenizer.stoi[\"<eos>\"]:\n        break\n```\nto\n```\nend_condition = torch.zeros(batch_size, dtype=torch.long).to(encoder_out.device)\nfor t in range(decode_lengths):\n   ...\n   end_condition |= (torch.argmax(preds, -1) == tokenizer.stoi[\"<eos>\"])\n       if end_condition.sum() == batch_size:\n           break\n```\n\nUsing torch.jit on the resnet34 encoder seems to have no effect.",
      "votes": 25
    },
    {
      "id": 1269888,
      "postDate": "2021-04-11T03:07:10.657Z",
      "content": "<p>if you are using vision transformer:</p>\n<p><a href=\"https://pytorch.org/tutorials/beginner/vt_tutorial.html\" target=\"_blank\">https://pytorch.org/tutorials/beginner/vt_tutorial.html</a></p>\n<pre><code>original model: 169.95ms\nscripted model: 134.65ms\nscripted &amp; quantized model: 125.27ms\nscripted &amp; quantized &amp; optimized model: 120.05ms\nlite model: 113.12ms\n</code></pre>",
      "rawMarkdown": "if you are using vision transformer:\n\nhttps://pytorch.org/tutorials/beginner/vt_tutorial.html\n```\n\noriginal model: 169.95ms\nscripted model: 134.65ms\nscripted & quantized model: 125.27ms\nscripted & quantized & optimized model: 120.05ms\nlite model: 113.12ms\n\n```\n",
      "votes": 3
    },
    {
      "id": 1267490,
      "postDate": "2021-04-08T14:54:42.703Z",
      "content": "<p>thanks for sharing，I don’t understand why sort by length can speed up inference? Could you give me some explanation about that, thanks.</p>",
      "rawMarkdown": "thanks for sharing，I don’t understand why sort by length can speed up inference? Could you give me some explanation about that, thanks.",
      "votes": 1,
      "replies": [
        {
          "id": 1267503,
          "postDate": "2021-04-08T15:04:18.183Z",
          "content": "<p>If you sort the data, the elements in your batch will have similar lengths and reach their EOS in similar time steps.</p>",
          "rawMarkdown": "If you sort the data, the elements in your batch will have similar lengths and reach their EOS in similar time steps.",
          "votes": 5
        },
        {
          "id": 1267517,
          "postDate": "2021-04-08T15:11:59.070Z",
          "content": "<p>Ok, I’ll try, thanks</p>",
          "rawMarkdown": "Ok, I’ll try, thanks"
        }
      ]
    },
    {
      "id": 1268260,
      "postDate": "2021-04-09T08:45:01.187Z",
      "content": "<p>i recall someone use this trick before:</p>\n<ol>\n<li>encode the image into embedding, e.g. resnet34</li>\n<li>use the embedding for seq model 1, 2,3 ….</li>\n<li>by sharing the image embedding we saved time in the feature extraction time</li>\n<li>we can also do an ensemble k-beam search</li>\n</ol>",
      "rawMarkdown": "i recall someone use this trick before:\n\n1. encode the image into embedding, e.g. resnet34\n2. use the embedding for seq model 1, 2,3 ....\n3. by sharing the image embedding we saved time in the feature extraction time\n4. we can also do an ensemble k-beam search",
      "votes": 2
    },
    {
      "id": 1267952,
      "postDate": "2021-04-09T01:58:47.347Z",
      "content": "<p>for those interested in jit, refer to thsi: <a href=\"https://pytorch.org/tutorials/beginner/deploy_seq2seq_hybrid_frontend_tutorial.html\" target=\"_blank\">https://pytorch.org/tutorials/beginner/deploy_seq2seq_hybrid_frontend_tutorial.html</a></p>",
      "rawMarkdown": "for those interested in jit, refer to thsi: https://pytorch.org/tutorials/beginner/deploy_seq2seq_hybrid_frontend_tutorial.html",
      "votes": 2,
      "replies": [
        {
          "id": 1268610,
          "postDate": "2021-04-09T15:27:39.730Z",
          "content": "<p>Not sure if this is the same…. but for those interested in JIT within the context of TF and XLA acceleration please refer to this:</p>\n<ul>\n<li><a href=\"https://www.tensorflow.org/api_docs/python/tf/config/optimizer/set_jit?version=nightly\" target=\"_blank\"><strong>JIT TF Documentation</strong></a></li>\n<li><a href=\"https://www.tensorflow.org/xla/tutorials/autoclustering_xla\" target=\"_blank\"><strong>XLA Optimization Tutorial Which Utilizes JIT</strong></a></li>\n</ul>",
          "rawMarkdown": "Not sure if this is the same.... but for those interested in JIT within the context of TF and XLA acceleration please refer to this:\n- [**JIT TF Documentation**](https://www.tensorflow.org/api_docs/python/tf/config/optimizer/set_jit?version=nightly)\n- [**XLA Optimization Tutorial Which Utilizes JIT**](https://www.tensorflow.org/xla/tutorials/autoclustering_xla)"
        }
      ]
    },
    {
      "id": 1267978,
      "postDate": "2021-04-09T03:06:52.720Z",
      "content": "<p>How about using TPU? It takes about 1 Hr in my case. Too long?</p>",
      "rawMarkdown": "How about using TPU? It takes about 1 Hr in my case. Too long?",
      "replies": [
        {
          "id": 1268159,
          "postDate": "2021-04-09T07:12:00.900Z",
          "content": "<p><a href=\"https://www.kaggle.com/wuliaokaola\" target=\"_blank\">@wuliaokaola</a> are you using TPUs for inference?<br>\n(I thought Kaggle allows using TPUs for training only, not inference.)</p>",
          "rawMarkdown": "@wuliaokaola are you using TPUs for inference?\n(I thought Kaggle allows using TPUs for training only, not inference.)"
        },
        {
          "id": 1268252,
          "postDate": "2021-04-09T08:39:01.013Z",
          "content": "<p><a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> This is not code competition. You can use TPU as you like. <br>\nBTW, I think code competition does not support TPU just because it need internet connection.</p>",
          "rawMarkdown": "@sirishks This is not code competition. You can use TPU as you like. \nBTW, I think code competition does not support TPU just because it need internet connection.",
          "votes": 1
        },
        {
          "id": 1268416,
          "postDate": "2021-04-09T11:28:40.287Z",
          "content": "<p>I’m using TPU. </p>\n<p>It takes 8-10 minutes for inference on 1.5 million images with an EfficientNetB3 and 384x384 images.</p>\n<p>This is the fastest I’ve seen. I’ll be publishing a notebook in a week or so showing my process for both training and testing using TPU in detail.</p>",
          "rawMarkdown": "I’m using TPU. \n\nIt takes 8-10 minutes for inference on 1.5 million images with an EfficientNetB3 and 384x384 images.\n\nThis is the fastest I’ve seen. I’ll be publishing a notebook in a week or so showing my process for both training and testing using TPU in detail."
        },
        {
          "id": 1268488,
          "postDate": "2021-04-09T12:59:46.090Z",
          "content": "<p>You mean TPU + TF with well sharded TFRecords ? </p>\n<p>Because you still have dataloader bottleneck with Pytorch/XLA which uses only local VM</p>",
          "rawMarkdown": "You mean TPU + TF with well sharded TFRecords ? \n\nBecause you still have dataloader bottleneck with Pytorch/XLA which uses only local VM\n"
        },
        {
          "id": 1268608,
          "postDate": "2021-04-09T15:25:59.940Z",
          "content": "<blockquote>\n  <p>TPU+TF with sharded TFRecords</p>\n</blockquote>\n<p>Yep! I'm not sorting the data or doing any tricks like that… so I will hopefully be able to get it even faster.</p>",
          "rawMarkdown": "> TPU+TF with sharded TFRecords\n\nYep! I'm not sorting the data or doing any tricks like that... so I will hopefully be able to get it even faster.",
          "votes": 1
        },
        {
          "id": 1270660,
          "postDate": "2021-04-11T21:08:17.993Z",
          "content": "<p>I tried to use distributed TPU ,however I  got  shared memory expired!  error. can not find any distributed TPU inference examples.</p>",
          "rawMarkdown": "I tried to use distributed TPU ,however I  got  shared memory expired!  error. can not find any distributed TPU inference examples."
        },
        {
          "id": 1279058,
          "postDate": "2021-04-20T14:53:37.117Z",
          "content": "<p>Here is my inference notebook - <a href=\"https://www.kaggle.com/dschettler8845/bms-cnn-attn-lstm-tpu-tutorial-infer\" target=\"_blank\">https://www.kaggle.com/dschettler8845/bms-cnn-attn-lstm-tpu-tutorial-infer</a></p>",
          "rawMarkdown": "Here is my inference notebook - https://www.kaggle.com/dschettler8845/bms-cnn-attn-lstm-tpu-tutorial-infer",
          "votes": 1
        },
        {
          "id": 1279181,
          "postDate": "2021-04-20T17:07:52.777Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1274092,
      "postDate": "2021-04-15T01:15:55.850Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1269888,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-04-11T03:07:10.657000",
      "content": "<p>if you are using vision transformer:</p>\n<p><a href=\"https://pytorch.org/tutorials/beginner/vt_tutorial.html\" target=\"_blank\">https://pytorch.org/tutorials/beginner/vt_tutorial.html</a></p>\n<pre><code>original model: 169.95ms\nscripted model: 134.65ms\nscripted &amp; quantized model: 125.27ms\nscripted &amp; quantized &amp; optimized model: 120.05ms\nlite model: 113.12ms\n</code></pre>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1267490,
      "author_name": "Matrix",
      "author_url": "",
      "post_date": "2021-04-08T14:54:42.703000",
      "content": "<p>thanks for sharing，I don’t understand why sort by length can speed up inference? Could you give me some explanation about that, thanks.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1267503,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "2021-04-08T15:04:18.183000",
          "content": "<p>If you sort the data, the elements in your batch will have similar lengths and reach their EOS in similar time steps.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1267517,
          "author_name": "Matrix",
          "author_url": "",
          "post_date": "2021-04-08T15:11:59.070000",
          "content": "<p>Ok, I’ll try, thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1268260,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-04-09T08:45:01.187000",
      "content": "<p>i recall someone use this trick before:</p>\n<ol>\n<li>encode the image into embedding, e.g. resnet34</li>\n<li>use the embedding for seq model 1, 2,3 ….</li>\n<li>by sharing the image embedding we saved time in the feature extraction time</li>\n<li>we can also do an ensemble k-beam search</li>\n</ol>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1267952,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-04-09T01:58:47.347000",
      "content": "<p>for those interested in jit, refer to thsi: <a href=\"https://pytorch.org/tutorials/beginner/deploy_seq2seq_hybrid_frontend_tutorial.html\" target=\"_blank\">https://pytorch.org/tutorials/beginner/deploy_seq2seq_hybrid_frontend_tutorial.html</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 1268610,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-04-09T15:27:39.730000",
          "content": "<p>Not sure if this is the same…. but for those interested in JIT within the context of TF and XLA acceleration please refer to this:</p>\n<ul>\n<li><a href=\"https://www.tensorflow.org/api_docs/python/tf/config/optimizer/set_jit?version=nightly\" target=\"_blank\"><strong>JIT TF Documentation</strong></a></li>\n<li><a href=\"https://www.tensorflow.org/xla/tutorials/autoclustering_xla\" target=\"_blank\"><strong>XLA Optimization Tutorial Which Utilizes JIT</strong></a></li>\n</ul>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1267978,
      "author_name": "Johnny Lee",
      "author_url": "",
      "post_date": "2021-04-09T03:06:52.720000",
      "content": "<p>How about using TPU? It takes about 1 Hr in my case. Too long?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1268159,
          "author_name": "Sirish Somanchi",
          "author_url": "",
          "post_date": "2021-04-09T07:12:00.900000",
          "content": "<p><a href=\"https://www.kaggle.com/wuliaokaola\" target=\"_blank\">@wuliaokaola</a> are you using TPUs for inference?<br>\n(I thought Kaggle allows using TPUs for training only, not inference.)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1268252,
          "author_name": "Johnny Lee",
          "author_url": "",
          "post_date": "2021-04-09T08:39:01.013000",
          "content": "<p><a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> This is not code competition. You can use TPU as you like. <br>\nBTW, I think code competition does not support TPU just because it need internet connection.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1268416,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-04-09T11:28:40.287000",
          "content": "<p>I’m using TPU. </p>\n<p>It takes 8-10 minutes for inference on 1.5 million images with an EfficientNetB3 and 384x384 images.</p>\n<p>This is the fastest I’ve seen. I’ll be publishing a notebook in a week or so showing my process for both training and testing using TPU in detail.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1268488,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2021-04-09T12:59:46.090000",
          "content": "<p>You mean TPU + TF with well sharded TFRecords ? </p>\n<p>Because you still have dataloader bottleneck with Pytorch/XLA which uses only local VM</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1268608,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-04-09T15:25:59.940000",
          "content": "<blockquote>\n  <p>TPU+TF with sharded TFRecords</p>\n</blockquote>\n<p>Yep! I'm not sorting the data or doing any tricks like that… so I will hopefully be able to get it even faster.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1270660,
          "author_name": "dragon zhang",
          "author_url": "",
          "post_date": "2021-04-11T21:08:17.993000",
          "content": "<p>I tried to use distributed TPU ,however I  got  shared memory expired!  error. can not find any distributed TPU inference examples.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1279058,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-04-20T14:53:37.117000",
          "content": "<p>Here is my inference notebook - <a href=\"https://www.kaggle.com/dschettler8845/bms-cnn-attn-lstm-tpu-tutorial-infer\" target=\"_blank\">https://www.kaggle.com/dschettler8845/bms-cnn-attn-lstm-tpu-tutorial-infer</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1279181,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-20T17:07:52.777000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1274092,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-15T01:15:55.850000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1267215": "Inferencing on 1.5 million images is really slow. So i made some experiments with resnet34 to accelerate the inferencing. Here are results on the validation set (20% of the train set):\n```\nGPU: 1x 2080ti\nencoder: resnet34\ndecoder: lstm with attention\nbatch size: 256\nworker: 8\nimage size: 224x224\n\n20m14s: inferencing like Y.Nakamas public notebook\n16m50s: break if batch EOS\n12m39s: break if batch EOS + sort by InChi length\n12m42s: break if batch EOS + sort by InChi length + torch.jit.script(encoder)\n```\n\nSo with sorting by InChi length, you can make your inferencing 35% faster if you have already a previous submission. If not, you can use batch EOS break which is still much faster than the naive inferencing.\n\nTo use the batch EOS break, change the Y.Nakamas code this way:\n```\nfor t in range(decode_lengths):\n    ...\n    if np.argmax(preds.detach().cpu().numpy()) == tokenizer.stoi[\"<eos>\"]:\n        break\n```\nto\n```\nend_condition = torch.zeros(batch_size, dtype=torch.long).to(encoder_out.device)\nfor t in range(decode_lengths):\n   ...\n   end_condition |= (torch.argmax(preds, -1) == tokenizer.stoi[\"<eos>\"])\n       if end_condition.sum() == batch_size:\n           break\n```\n\nUsing torch.jit on the resnet34 encoder seems to have no effect.",
    "1269888": "if you are using vision transformer:\n\nhttps://pytorch.org/tutorials/beginner/vt_tutorial.html\n```\n\noriginal model: 169.95ms\nscripted model: 134.65ms\nscripted & quantized model: 125.27ms\nscripted & quantized & optimized model: 120.05ms\nlite model: 113.12ms\n\n```\n",
    "1267490": "thanks for sharing，I don’t understand why sort by length can speed up inference? Could you give me some explanation about that, thanks.",
    "1268260": "i recall someone use this trick before:\n\n1. encode the image into embedding, e.g. resnet34\n2. use the embedding for seq model 1, 2,3 ....\n3. by sharing the image embedding we saved time in the feature extraction time\n4. we can also do an ensemble k-beam search",
    "1267952": "for those interested in jit, refer to thsi: https://pytorch.org/tutorials/beginner/deploy_seq2seq_hybrid_frontend_tutorial.html",
    "1267978": "How about using TPU? It takes about 1 Hr in my case. Too long?",
    "1274092": ""
  }
}