{
  "id": 228975,
  "title": "how to improve inference speed??",
  "url": "/competitions/bms-molecular-translation/discussion/228975",
  "author_name": "",
  "post_date": "2021-03-27T14:28:53.848484200Z",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>HI, <br>\nmy model is taking approximately around 1sec for each sample to infer. it takes ages to complete 1.6m. i have used with tf.device(gpu:0) but not much improvement. batching also. <br>\nany suggestion to improve inference speed? </p>",
  "messages": [
    {
      "id": "1254321",
      "postDate": "03/27/2021 14:28:53",
      "content": "<p>HI, <br>\nmy model is taking approximately around 1sec for each sample to infer. it takes ages to complete 1.6m. i have used with tf.device(gpu:0) but not much improvement. batching also. <br>\nany suggestion to improve inference speed? </p>",
      "rawMarkdown": "HI, \nmy model is taking approximately around 1sec for each sample to infer. it takes ages to complete 1.6m. i have used with tf.device(gpu:0) but not much improvement. batching also. \nany suggestion to improve inference speed?",
      "votes": null
    },
    {
      "id": "1254401",
      "postDate": "03/27/2021 16:38:29",
      "content": "<p>If it’s autoregressive maybe you’d want to break if an EOS token appears. How to batch it? Same concept but break when every output sequence in the batch has an EOS token. </p>",
      "rawMarkdown": "If it’s autoregressive maybe you’d want to break if an EOS token appears. How to batch it? Same concept but break when every output sequence in the batch has an EOS token.",
      "votes": null
    },
    {
      "id": "1254423",
      "postDate": "03/27/2021 16:58:02",
      "content": "<p>Exactly! Also keep track of where in each sample of the batch you encountered the EOS first.</p>",
      "rawMarkdown": "Exactly! Also keep track of where in each sample of the batch you encountered the EOS first.",
      "votes": null
    },
    {
      "id": "1262502",
      "postDate": "04/04/2021 11:39:41",
      "content": "<p>After you have a decent submission, you can sort the test samples by output length and batch them according to that the next time you create your submission. This way you have almost same amount of prediction steps per sample in a batch and you don't waste much time on computing on samples that already finished.</p>\n<p>If I'm not mistaken, this should make it about x2 faster.</p>",
      "rawMarkdown": "After you have a decent submission, you can sort the test samples by output length and batch them according to that the next time you create your submission. This way you have almost same amount of prediction steps per sample in a batch and you don't waste much time on computing on samples that already finished.\n\nIf I'm not mistaken, this should make it about x2 faster.",
      "votes": null
    },
    {
      "id": "1279064",
      "postDate": "04/20/2021 14:57:21",
      "content": "<p>I prefer to use larger batch sizes (get my parallelization there) and just infer up to max length.</p>\n<p>I can do inference on the entire test dataset in 5-15 minutes (EfficientNet models and 512/1024 Attention/LSTM).</p>\n<p>Here is the <a href=\"https://www.kaggle.com/dschettler8845/bms-cnn-attn-lstm-tpu-tutorial-infer\" target=\"_blank\"><strong>inference notebook</strong></a> if you are curious - <a href=\"https://www.kaggle.com/dschettler8845/bms-cnn-attn-lstm-tpu-tutorial-infer\" target=\"_blank\">https://www.kaggle.com/dschettler8845/bms-cnn-attn-lstm-tpu-tutorial-infer</a></p>\n<p>This is with a EfficientNetB3 and I think it takes around 8-9 minutes for inference and around 8 minutes to post-process (convert arrays into inchi strings… could probably speed this up).</p>\n<hr>\n<p>ps: The 15 minute inference time is what gives me my current Public LB (~5.7) w/ an EfficientNETB7.</p>",
      "rawMarkdown": "I prefer to use larger batch sizes (get my parallelization there) and just infer up to max length.\n\nI can do inference on the entire test dataset in 5-15 minutes (EfficientNet models and 512/1024 Attention/LSTM).\n\nHere is the [**inference notebook**](https://www.kaggle.com/dschettler8845/bms-cnn-attn-lstm-tpu-tutorial-infer) if you are curious - https://www.kaggle.com/dschettler8845/bms-cnn-attn-lstm-tpu-tutorial-infer\n\nThis is with a EfficientNetB3 and I think it takes around 8-9 minutes for inference and around 8 minutes to post-process (convert arrays into inchi strings... could probably speed this up).\n\n---\n\nps: The 15 minute inference time is what gives me my current Public LB (~5.7) w/ an EfficientNETB7.",
      "votes": null
    },
    {
      "id": "1279169",
      "postDate": "04/20/2021 16:53:15",
      "content": "<p><code>inference on the entire test dataset in 5-15 minutes</code><br>\nThat seems to be insanely fast. But on the other hand you only have one sequential layer in your decoder. Maybe I should think about increasing the batch size…<br>\nI also think now, that my transformer decoder with 41M parameters might be way overkill. Although I don't surpass validation performance with train that fast. But faster then with enwik8 benchmark at least. Hmmm</p>\n<p>Your explanations in the notebook are how everyone should do it. Take my upvote for that. <br>\n(Although I use Pytorch and prefer to write my own code.)</p>",
      "rawMarkdown": "` inference on the entire test dataset in 5-15 minutes`\nThat seems to be insanely fast. But on the other hand you only have one sequential layer in your decoder. Maybe I should think about increasing the batch size...\nI also think now, that my transformer decoder with 41M parameters might be way overkill. Although I don't surpass validation performance with train that fast. But faster then with enwik8 benchmark at least. Hmmm\n\nYour explanations in the notebook are how everyone should do it. Take my upvote for that. \n(Although I use Pytorch and prefer to write my own code.)",
      "votes": null
    },
    {
      "id": "1310982",
      "postDate": "05/17/2021 05:02:50",
      "content": "<p>When predict is called with 20,000 inputs and batch_size=512, with Kaggle's GPU environment it takes 4 hours to process 1.6M test samples (2098 encoded image features + LSTM inputs + greedy algorithm).<br>\nIn my case, all the images are encoded prior to the inference.</p>",
      "rawMarkdown": "When predict is called with 20,000 inputs and batch_size=512, with Kaggle's GPU environment it takes 4 hours to process 1.6M test samples (2098 encoded image features + LSTM inputs + greedy algorithm).\nIn my case, all the images are encoded prior to the inference.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1254401,
      "author_name": "claverru",
      "author_url": "",
      "post_date": "03/27/2021 16:38:29",
      "content": "<p>If it’s autoregressive maybe you’d want to break if an EOS token appears. How to batch it? Same concept but break when every output sequence in the batch has an EOS token. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1254423,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "03/27/2021 16:58:02",
          "content": "<p>Exactly! Also keep track of where in each sample of the batch you encountered the EOS first.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1262502,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "04/04/2021 11:39:41",
          "content": "<p>After you have a decent submission, you can sort the test samples by output length and batch them according to that the next time you create your submission. This way you have almost same amount of prediction steps per sample in a batch and you don't waste much time on computing on samples that already finished.</p>\n<p>If I'm not mistaken, this should make it about x2 faster.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1279064,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "04/20/2021 14:57:21",
          "content": "<p>I prefer to use larger batch sizes (get my parallelization there) and just infer up to max length.</p>\n<p>I can do inference on the entire test dataset in 5-15 minutes (EfficientNet models and 512/1024 Attention/LSTM).</p>\n<p>Here is the <a href=\"https://www.kaggle.com/dschettler8845/bms-cnn-attn-lstm-tpu-tutorial-infer\" target=\"_blank\"><strong>inference notebook</strong></a> if you are curious - <a href=\"https://www.kaggle.com/dschettler8845/bms-cnn-attn-lstm-tpu-tutorial-infer\" target=\"_blank\">https://www.kaggle.com/dschettler8845/bms-cnn-attn-lstm-tpu-tutorial-infer</a></p>\n<p>This is with a EfficientNetB3 and I think it takes around 8-9 minutes for inference and around 8 minutes to post-process (convert arrays into inchi strings… could probably speed this up).</p>\n<hr>\n<p>ps: The 15 minute inference time is what gives me my current Public LB (~5.7) w/ an EfficientNETB7.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1279169,
          "author_name": "cepheidq",
          "author_url": "",
          "post_date": "04/20/2021 16:53:15",
          "content": "<p><code>inference on the entire test dataset in 5-15 minutes</code><br>\nThat seems to be insanely fast. But on the other hand you only have one sequential layer in your decoder. Maybe I should think about increasing the batch size…<br>\nI also think now, that my transformer decoder with 41M parameters might be way overkill. Although I don't surpass validation performance with train that fast. But faster then with enwik8 benchmark at least. Hmmm</p>\n<p>Your explanations in the notebook are how everyone should do it. Take my upvote for that. <br>\n(Although I use Pytorch and prefer to write my own code.)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1310982,
      "author_name": "jhasanov",
      "author_url": "",
      "post_date": "05/17/2021 05:02:50",
      "content": "<p>When predict is called with 20,000 inputs and batch_size=512, with Kaggle's GPU environment it takes 4 hours to process 1.6M test samples (2098 encoded image features + LSTM inputs + greedy algorithm).<br>\nIn my case, all the images are encoded prior to the inference.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1254321": "HI, \nmy model is taking approximately around 1sec for each sample to infer. it takes ages to complete 1.6m. i have used with tf.device(gpu:0) but not much improvement. batching also. \nany suggestion to improve inference speed?",
    "1254401": "If it’s autoregressive maybe you’d want to break if an EOS token appears. How to batch it? Same concept but break when every output sequence in the batch has an EOS token.",
    "1254423": "Exactly! Also keep track of where in each sample of the batch you encountered the EOS first.",
    "1262502": "After you have a decent submission, you can sort the test samples by output length and batch them according to that the next time you create your submission. This way you have almost same amount of prediction steps per sample in a batch and you don't waste much time on computing on samples that already finished.\n\nIf I'm not mistaken, this should make it about x2 faster.",
    "1279064": "I prefer to use larger batch sizes (get my parallelization there) and just infer up to max length.\n\nI can do inference on the entire test dataset in 5-15 minutes (EfficientNet models and 512/1024 Attention/LSTM).\n\nHere is the [**inference notebook**](https://www.kaggle.com/dschettler8845/bms-cnn-attn-lstm-tpu-tutorial-infer) if you are curious - https://www.kaggle.com/dschettler8845/bms-cnn-attn-lstm-tpu-tutorial-infer\n\nThis is with a EfficientNetB3 and I think it takes around 8-9 minutes for inference and around 8 minutes to post-process (convert arrays into inchi strings... could probably speed this up).\n\n---\n\nps: The 15 minute inference time is what gives me my current Public LB (~5.7) w/ an EfficientNETB7.",
    "1279169": "` inference on the entire test dataset in 5-15 minutes`\nThat seems to be insanely fast. But on the other hand you only have one sequential layer in your decoder. Maybe I should think about increasing the batch size...\nI also think now, that my transformer decoder with 41M parameters might be way overkill. Although I don't surpass validation performance with train that fast. But faster then with enwik8 benchmark at least. Hmmm\n\nYour explanations in the notebook are how everyone should do it. Take my upvote for that. \n(Although I use Pytorch and prefer to write my own code.)",
    "1310982": "When predict is called with 20,000 inputs and batch_size=512, with Kaggle's GPU environment it takes 4 hours to process 1.6M test samples (2098 encoded image features + LSTM inputs + greedy algorithm).\nIn my case, all the images are encoded prior to the inference."
  },
  "source": "meta"
}