{
  "id": 314204,
  "title": "Storytime –AKA– How I Painfully Learned About .predict vs. .__call__",
  "url": "/competitions/birdclef-2022/discussion/314204",
  "author_name": "",
  "post_date": "2022-03-21T14:02:01.663263900Z",
  "votes": 33,
  "comment_count": 2,
  "views": 0,
  "content": "<p>In the past, I've posted regarding my failures or mistakes and received very encouraging feedback that this type of post is helpful for others. So here we go.</p>\n<p><strong>The tl;dr</strong> </p>\n<blockquote>\n  <p><strong><code>model.predict()</code></strong> should only be used on datasets or items too large to fit in a single batch… otherwise it will cause a memory leak that grows and will eventually cause the RAM to overflow and crash your notebook. Instead you should use <strong><code>model.__call__()</code></strong> (i.e. <code>pred_batch=model(x_batch)</code> vs. <code>pred_batch=model.predict(x_batch)</code>).</p>\n</blockquote>\n<hr>\n<p><strong>The Long Story</strong></p>\n<p>I spent a WHILE building an incredibly rigorous pipeline for this competition (that I hope to open-source soon), with very clean code documentation and good separation of functionality between respective pieces.</p>\n<p>Everything was going very well and I had finished training my models and was ready for inference. I spent a day or so writing a notebook to create the desired output –&nbsp;not as easy as I had anticipated due to the difference in clip length between what my model expected (7-second clips) and what the prediction format required (5-second clips) –&nbsp;and I felt pretty confident that it would work.</p>\n<p><strong>So I submitted it… I waited 9 long hours before seeing it timeout.</strong></p>\n<p>Ok, so I realized I hadn't really validated the inference notebook at all… so I probably made a mistake. So I created a pseudo-test-set from the training dataset that was 1/5th of the size of the expected test set and formatted it similarly. I figured this would allow me to see how long each step in my pipeline took and would let me identify any errors.</p>\n<p><strong>I was right! I was accidentally preprocessing a single file 252 times (21*12=252) !! 😅😥😅</strong></p>\n<p>Ok, phew, found it. I then fixed that, ran through the validation version of my inference notebook and SUCCESS! </p>\n<p><strong>Now it was time to try to submit again and see what happens… I only had to wait 2 hours this time before seeing that the submission had thrown an Exception --&gt; Notebook Exceeded Allowed Compute</strong></p>\n<p>This is quite frustrating. Nothing in my validation notebook had in any way indicated that it was consuming even 1/10th of the possible compute (RAM, HD, CPU, GPU, etc.). That being said, I was dutiful and tried adjusting things to reduce the complexity of what I was doing. I tried</p>\n<ul>\n<li>Lowering the batch size</li>\n<li>Saving intermediate files in /kaggle/tmp instead of /kaggle/working</li>\n<li>Using a smaller sliding window (this reduced the number of inferences by 2/3rd in my pipeline)</li>\n<li>Set <strong><code>tf.config.experimental.set_memory_growth()</code></strong> properly</li>\n<li>Add in <strong><code>gc.collect()</code></strong> and <strong><code>tf.keras.backend.clear_session()</code></strong> where possible</li>\n</ul>\n<p><strong>But still I was getting the Notebook Exceeded Allowed Compute error!</strong></p>\n<p>I have two strategies left, although at this point I feel quite worn down. </p>\n<p><strong>Strategy 1</strong></p>\n<ul>\n<li>Rerun my validation inference notebook but make the dummy dataset EXACTLY the same size as the expected test dataset even though it feels unnecessary.</li>\n<li>This will allow me to see if any memory errors are thrown and target the cell in questions with more troubleshooting</li>\n</ul>\n<p><strong>Strategy 2</strong></p>\n<ul>\n<li>Change my inference submission so that instead of submitting MY submission at the end, I simply submit a dummy, all-False, submission… </li>\n<li>I thought maybe it was being caused by my submission violating one of the rules?</li>\n<li>I never had to use this strategy</li>\n</ul>\n<p>So I did Strategy 1… and lo and behold during the code block for model inference the notebook throws a memory exception and restarts… but HOW?! It wasn't even close before!? Let me tell you…</p>\n<p>**I was using *<em><code>model.predict</code></em>* on individual batches of data. (I actually have multiple models so I was calling it multiple times PER batch)**</p>\n<ul>\n<li>I did this because I had a tf.data.Dataset that contained <strong><code>row_id</code></strong> information and I wanted to ensure deterministic order… so I iterated through the batches storing the <strong><code>row_id</code></strong>s and respective predictions batch by batch (and then concat at the end).</li>\n</ul>\n<p>It turns out that calling <strong><code>model.predict</code></strong> results in a memory leak if called too frequently (this is kind of a <a href=\"https://github.com/keras-team/keras/issues/13118\" target=\"_blank\">known issue</a>… but the prevalence seems to come and go w/ different versions of TF). The documentation for the <strong><code>predict</code></strong> method even states this in the doc string:</p>\n<blockquote>\n  <p>For small numbers of inputs that fit in one batch,<br>\n      directly use <code>__call__()</code> for faster execution, e.g.,<br>\n      <code>model(x)</code>, or <code>model(x, training=False)</code> if you have layers such as<br>\n      <code>tf.keras.layers.BatchNormalization</code> that behave differently during<br>\n      inference. </p>\n</blockquote>\n<p>So… I guess that's the problem? So I decided to be extra safe. I changed the <strong><code>.predict()</code></strong> to a <strong><code>.__call__()</code></strong> method in my code and ran my full-length dummy inference validation notebook. On this run, it made it through the code block where it had previously failed and generated the required submission file. I then took this and submitted it and… finally… I was able to get a score!</p>\n<hr>\n<p><strong>The Conclusion &amp; Takeaways</strong></p>\n<p>That's my incredibly long-winded story about how I discovered the shortcoming/memory-leak found within the <strong><code>.predict</code></strong> method of a tf.Keras model. </p>\n<p>Hopefully, this shows you that we all make mistakes. We all waste lots of time trying things that don't work and getting frustrated as a result. Sometimes it doesn't matter how well you prepare or how organized you are when you approach things… there's always room to learn and improve. This is certainly the case for me and, while it was incredibly frustrating, I feel like I learned a lot from this experience.</p>\n<p>I hope you can learn from my mistakes and that you enjoyed (or at least empathized with) my frustration and plight 👹. Maybe next time you are struggling you'll know you are not alone, and that we've all been there. Keep trying. Fight through. Because the feeling of accomplishment is absolutely worth it when you crack through to the other side.</p>",
  "messages": [
    {
      "id": "1730659",
      "postDate": "03/21/2022 14:02:01",
      "content": "<p>In the past, I've posted regarding my failures or mistakes and received very encouraging feedback that this type of post is helpful for others. So here we go.</p>\n<p><strong>The tl;dr</strong> </p>\n<blockquote>\n  <p><strong><code>model.predict()</code></strong> should only be used on datasets or items too large to fit in a single batch… otherwise it will cause a memory leak that grows and will eventually cause the RAM to overflow and crash your notebook. Instead you should use <strong><code>model.__call__()</code></strong> (i.e. <code>pred_batch=model(x_batch)</code> vs. <code>pred_batch=model.predict(x_batch)</code>).</p>\n</blockquote>\n<hr>\n<p><strong>The Long Story</strong></p>\n<p>I spent a WHILE building an incredibly rigorous pipeline for this competition (that I hope to open-source soon), with very clean code documentation and good separation of functionality between respective pieces.</p>\n<p>Everything was going very well and I had finished training my models and was ready for inference. I spent a day or so writing a notebook to create the desired output –&nbsp;not as easy as I had anticipated due to the difference in clip length between what my model expected (7-second clips) and what the prediction format required (5-second clips) –&nbsp;and I felt pretty confident that it would work.</p>\n<p><strong>So I submitted it… I waited 9 long hours before seeing it timeout.</strong></p>\n<p>Ok, so I realized I hadn't really validated the inference notebook at all… so I probably made a mistake. So I created a pseudo-test-set from the training dataset that was 1/5th of the size of the expected test set and formatted it similarly. I figured this would allow me to see how long each step in my pipeline took and would let me identify any errors.</p>\n<p><strong>I was right! I was accidentally preprocessing a single file 252 times (21*12=252) !! 😅😥😅</strong></p>\n<p>Ok, phew, found it. I then fixed that, ran through the validation version of my inference notebook and SUCCESS! </p>\n<p><strong>Now it was time to try to submit again and see what happens… I only had to wait 2 hours this time before seeing that the submission had thrown an Exception --&gt; Notebook Exceeded Allowed Compute</strong></p>\n<p>This is quite frustrating. Nothing in my validation notebook had in any way indicated that it was consuming even 1/10th of the possible compute (RAM, HD, CPU, GPU, etc.). That being said, I was dutiful and tried adjusting things to reduce the complexity of what I was doing. I tried</p>\n<ul>\n<li>Lowering the batch size</li>\n<li>Saving intermediate files in /kaggle/tmp instead of /kaggle/working</li>\n<li>Using a smaller sliding window (this reduced the number of inferences by 2/3rd in my pipeline)</li>\n<li>Set <strong><code>tf.config.experimental.set_memory_growth()</code></strong> properly</li>\n<li>Add in <strong><code>gc.collect()</code></strong> and <strong><code>tf.keras.backend.clear_session()</code></strong> where possible</li>\n</ul>\n<p><strong>But still I was getting the Notebook Exceeded Allowed Compute error!</strong></p>\n<p>I have two strategies left, although at this point I feel quite worn down. </p>\n<p><strong>Strategy 1</strong></p>\n<ul>\n<li>Rerun my validation inference notebook but make the dummy dataset EXACTLY the same size as the expected test dataset even though it feels unnecessary.</li>\n<li>This will allow me to see if any memory errors are thrown and target the cell in questions with more troubleshooting</li>\n</ul>\n<p><strong>Strategy 2</strong></p>\n<ul>\n<li>Change my inference submission so that instead of submitting MY submission at the end, I simply submit a dummy, all-False, submission… </li>\n<li>I thought maybe it was being caused by my submission violating one of the rules?</li>\n<li>I never had to use this strategy</li>\n</ul>\n<p>So I did Strategy 1… and lo and behold during the code block for model inference the notebook throws a memory exception and restarts… but HOW?! It wasn't even close before!? Let me tell you…</p>\n<p>**I was using *<em><code>model.predict</code></em>* on individual batches of data. (I actually have multiple models so I was calling it multiple times PER batch)**</p>\n<ul>\n<li>I did this because I had a tf.data.Dataset that contained <strong><code>row_id</code></strong> information and I wanted to ensure deterministic order… so I iterated through the batches storing the <strong><code>row_id</code></strong>s and respective predictions batch by batch (and then concat at the end).</li>\n</ul>\n<p>It turns out that calling <strong><code>model.predict</code></strong> results in a memory leak if called too frequently (this is kind of a <a href=\"https://github.com/keras-team/keras/issues/13118\" target=\"_blank\">known issue</a>… but the prevalence seems to come and go w/ different versions of TF). The documentation for the <strong><code>predict</code></strong> method even states this in the doc string:</p>\n<blockquote>\n  <p>For small numbers of inputs that fit in one batch,<br>\n      directly use <code>__call__()</code> for faster execution, e.g.,<br>\n      <code>model(x)</code>, or <code>model(x, training=False)</code> if you have layers such as<br>\n      <code>tf.keras.layers.BatchNormalization</code> that behave differently during<br>\n      inference. </p>\n</blockquote>\n<p>So… I guess that's the problem? So I decided to be extra safe. I changed the <strong><code>.predict()</code></strong> to a <strong><code>.__call__()</code></strong> method in my code and ran my full-length dummy inference validation notebook. On this run, it made it through the code block where it had previously failed and generated the required submission file. I then took this and submitted it and… finally… I was able to get a score!</p>\n<hr>\n<p><strong>The Conclusion &amp; Takeaways</strong></p>\n<p>That's my incredibly long-winded story about how I discovered the shortcoming/memory-leak found within the <strong><code>.predict</code></strong> method of a tf.Keras model. </p>\n<p>Hopefully, this shows you that we all make mistakes. We all waste lots of time trying things that don't work and getting frustrated as a result. Sometimes it doesn't matter how well you prepare or how organized you are when you approach things… there's always room to learn and improve. This is certainly the case for me and, while it was incredibly frustrating, I feel like I learned a lot from this experience.</p>\n<p>I hope you can learn from my mistakes and that you enjoyed (or at least empathized with) my frustration and plight 👹. Maybe next time you are struggling you'll know you are not alone, and that we've all been there. Keep trying. Fight through. Because the feeling of accomplishment is absolutely worth it when you crack through to the other side.</p>",
      "rawMarkdown": "In the past, I've posted regarding my failures or mistakes and received very encouraging feedback that this type of post is helpful for others. So here we go.\n\n**The tl;dr** \n\n> **`model.predict()`** should only be used on datasets or items too large to fit in a single batch... otherwise it will cause a memory leak that grows and will eventually cause the RAM to overflow and crash your notebook. Instead you should use **`model.__call__()`** (i.e. `pred_batch=model(x_batch)` vs. `pred_batch=model.predict(x_batch)`).\n\n---\n\n**The Long Story**\n\nI spent a WHILE building an incredibly rigorous pipeline for this competition (that I hope to open-source soon), with very clean code documentation and good separation of functionality between respective pieces.\n\nEverything was going very well and I had finished training my models and was ready for inference. I spent a day or so writing a notebook to create the desired output – not as easy as I had anticipated due to the difference in clip length between what my model expected (7-second clips) and what the prediction format required (5-second clips) – and I felt pretty confident that it would work.\n\n**So I submitted it... I waited 9 long hours before seeing it timeout.**\n\nOk, so I realized I hadn't really validated the inference notebook at all... so I probably made a mistake. So I created a pseudo-test-set from the training dataset that was 1/5th of the size of the expected test set and formatted it similarly. I figured this would allow me to see how long each step in my pipeline took and would let me identify any errors.\n\n**I was right! I was accidentally preprocessing a single file 252 times (21*12=252) !! 😅😥😅**\n\nOk, phew, found it. I then fixed that, ran through the validation version of my inference notebook and SUCCESS! \n\n**Now it was time to try to submit again and see what happens... I only had to wait 2 hours this time before seeing that the submission had thrown an Exception --> Notebook Exceeded Allowed Compute**\n\nThis is quite frustrating. Nothing in my validation notebook had in any way indicated that it was consuming even 1/10th of the possible compute (RAM, HD, CPU, GPU, etc.). That being said, I was dutiful and tried adjusting things to reduce the complexity of what I was doing. I tried\n* Lowering the batch size\n* Saving intermediate files in /kaggle/tmp instead of /kaggle/working\n* Using a smaller sliding window (this reduced the number of inferences by 2/3rd in my pipeline)\n* Set **`tf.config.experimental.set_memory_growth()`** properly\n* Add in **`gc.collect()`** and **`tf.keras.backend.clear_session()`** where possible\n\n**But still I was getting the Notebook Exceeded Allowed Compute error!**\n\nI have two strategies left, although at this point I feel quite worn down. \n\n**Strategy 1**\n* Rerun my validation inference notebook but make the dummy dataset EXACTLY the same size as the expected test dataset even though it feels unnecessary.\n* This will allow me to see if any memory errors are thrown and target the cell in questions with more troubleshooting\n\n**Strategy 2**\n* Change my inference submission so that instead of submitting MY submission at the end, I simply submit a dummy, all-False, submission... \n* I thought maybe it was being caused by my submission violating one of the rules?\n* I never had to use this strategy\n\nSo I did Strategy 1... and lo and behold during the code block for model inference the notebook throws a memory exception and restarts... but HOW?! It wasn't even close before!? Let me tell you...\n\n**I was using **`model.predict`** on individual batches of data. (I actually have multiple models so I was calling it multiple times PER batch)**\n* I did this because I had a tf.data.Dataset that contained **`row_id`** information and I wanted to ensure deterministic order... so I iterated through the batches storing the **`row_id`**s and respective predictions batch by batch (and then concat at the end).\n\nIt turns out that calling **`model.predict`** results in a memory leak if called too frequently (this is kind of a [known issue](https://github.com/keras-team/keras/issues/13118)... but the prevalence seems to come and go w/ different versions of TF). The documentation for the **`predict`** method even states this in the doc string:\n\n> For small numbers of inputs that fit in one batch,\n    directly use `__call__()` for faster execution, e.g.,\n    `model(x)`, or `model(x, training=False)` if you have layers such as\n    `tf.keras.layers.BatchNormalization` that behave differently during\n    inference. \n\nSo... I guess that's the problem? So I decided to be extra safe. I changed the **`.predict()`** to a **`.__call__()`** method in my code and ran my full-length dummy inference validation notebook. On this run, it made it through the code block where it had previously failed and generated the required submission file. I then took this and submitted it and... finally... I was able to get a score!\n\n---\n\n**The Conclusion & Takeaways**\n\nThat's my incredibly long-winded story about how I discovered the shortcoming/memory-leak found within the **`.predict`** method of a tf.Keras model. \n\nHopefully, this shows you that we all make mistakes. We all waste lots of time trying things that don't work and getting frustrated as a result. Sometimes it doesn't matter how well you prepare or how organized you are when you approach things... there's always room to learn and improve. This is certainly the case for me and, while it was incredibly frustrating, I feel like I learned a lot from this experience.\n\nI hope you can learn from my mistakes and that you enjoyed (or at least empathized with) my frustration and plight 👹. Maybe next time you are struggling you'll know you are not alone, and that we've all been there. Keep trying. Fight through. Because the feeling of accomplishment is absolutely worth it when you crack through to the other side.",
      "votes": null
    },
    {
      "id": "1730761",
      "postDate": "03/21/2022 15:33:03",
      "content": "<p>You've been through a lot.<br>\nMemory errors in code competition really annoy me.</p>",
      "rawMarkdown": "You've been through a lot.\nMemory errors in code competition really annoy me.",
      "votes": null
    },
    {
      "id": "1799664",
      "postDate": "05/24/2022 07:11:30",
      "content": "<p>Very helpful! Thank you for sharing!</p>",
      "rawMarkdown": "Very helpful! Thank you for sharing!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1730761,
      "author_name": "gwanghan",
      "author_url": "",
      "post_date": "03/21/2022 15:33:03",
      "content": "<p>You've been through a lot.<br>\nMemory errors in code competition really annoy me.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1799664,
      "author_name": "jamesbarthelemy",
      "author_url": "",
      "post_date": "05/24/2022 07:11:30",
      "content": "<p>Very helpful! Thank you for sharing!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1730659": "In the past, I've posted regarding my failures or mistakes and received very encouraging feedback that this type of post is helpful for others. So here we go.\n\n**The tl;dr** \n\n> **`model.predict()`** should only be used on datasets or items too large to fit in a single batch... otherwise it will cause a memory leak that grows and will eventually cause the RAM to overflow and crash your notebook. Instead you should use **`model.__call__()`** (i.e. `pred_batch=model(x_batch)` vs. `pred_batch=model.predict(x_batch)`).\n\n---\n\n**The Long Story**\n\nI spent a WHILE building an incredibly rigorous pipeline for this competition (that I hope to open-source soon), with very clean code documentation and good separation of functionality between respective pieces.\n\nEverything was going very well and I had finished training my models and was ready for inference. I spent a day or so writing a notebook to create the desired output – not as easy as I had anticipated due to the difference in clip length between what my model expected (7-second clips) and what the prediction format required (5-second clips) – and I felt pretty confident that it would work.\n\n**So I submitted it... I waited 9 long hours before seeing it timeout.**\n\nOk, so I realized I hadn't really validated the inference notebook at all... so I probably made a mistake. So I created a pseudo-test-set from the training dataset that was 1/5th of the size of the expected test set and formatted it similarly. I figured this would allow me to see how long each step in my pipeline took and would let me identify any errors.\n\n**I was right! I was accidentally preprocessing a single file 252 times (21*12=252) !! 😅😥😅**\n\nOk, phew, found it. I then fixed that, ran through the validation version of my inference notebook and SUCCESS! \n\n**Now it was time to try to submit again and see what happens... I only had to wait 2 hours this time before seeing that the submission had thrown an Exception --> Notebook Exceeded Allowed Compute**\n\nThis is quite frustrating. Nothing in my validation notebook had in any way indicated that it was consuming even 1/10th of the possible compute (RAM, HD, CPU, GPU, etc.). That being said, I was dutiful and tried adjusting things to reduce the complexity of what I was doing. I tried\n* Lowering the batch size\n* Saving intermediate files in /kaggle/tmp instead of /kaggle/working\n* Using a smaller sliding window (this reduced the number of inferences by 2/3rd in my pipeline)\n* Set **`tf.config.experimental.set_memory_growth()`** properly\n* Add in **`gc.collect()`** and **`tf.keras.backend.clear_session()`** where possible\n\n**But still I was getting the Notebook Exceeded Allowed Compute error!**\n\nI have two strategies left, although at this point I feel quite worn down. \n\n**Strategy 1**\n* Rerun my validation inference notebook but make the dummy dataset EXACTLY the same size as the expected test dataset even though it feels unnecessary.\n* This will allow me to see if any memory errors are thrown and target the cell in questions with more troubleshooting\n\n**Strategy 2**\n* Change my inference submission so that instead of submitting MY submission at the end, I simply submit a dummy, all-False, submission... \n* I thought maybe it was being caused by my submission violating one of the rules?\n* I never had to use this strategy\n\nSo I did Strategy 1... and lo and behold during the code block for model inference the notebook throws a memory exception and restarts... but HOW?! It wasn't even close before!? Let me tell you...\n\n**I was using **`model.predict`** on individual batches of data. (I actually have multiple models so I was calling it multiple times PER batch)**\n* I did this because I had a tf.data.Dataset that contained **`row_id`** information and I wanted to ensure deterministic order... so I iterated through the batches storing the **`row_id`**s and respective predictions batch by batch (and then concat at the end).\n\nIt turns out that calling **`model.predict`** results in a memory leak if called too frequently (this is kind of a [known issue](https://github.com/keras-team/keras/issues/13118)... but the prevalence seems to come and go w/ different versions of TF). The documentation for the **`predict`** method even states this in the doc string:\n\n> For small numbers of inputs that fit in one batch,\n    directly use `__call__()` for faster execution, e.g.,\n    `model(x)`, or `model(x, training=False)` if you have layers such as\n    `tf.keras.layers.BatchNormalization` that behave differently during\n    inference. \n\nSo... I guess that's the problem? So I decided to be extra safe. I changed the **`.predict()`** to a **`.__call__()`** method in my code and ran my full-length dummy inference validation notebook. On this run, it made it through the code block where it had previously failed and generated the required submission file. I then took this and submitted it and... finally... I was able to get a score!\n\n---\n\n**The Conclusion & Takeaways**\n\nThat's my incredibly long-winded story about how I discovered the shortcoming/memory-leak found within the **`.predict`** method of a tf.Keras model. \n\nHopefully, this shows you that we all make mistakes. We all waste lots of time trying things that don't work and getting frustrated as a result. Sometimes it doesn't matter how well you prepare or how organized you are when you approach things... there's always room to learn and improve. This is certainly the case for me and, while it was incredibly frustrating, I feel like I learned a lot from this experience.\n\nI hope you can learn from my mistakes and that you enjoyed (or at least empathized with) my frustration and plight 👹. Maybe next time you are struggling you'll know you are not alone, and that we've all been there. Keep trying. Fight through. Because the feeling of accomplishment is absolutely worth it when you crack through to the other side.",
    "1730761": "You've been through a lot.\nMemory errors in code competition really annoy me.",
    "1799664": "Very helpful! Thank you for sharing!"
  },
  "source": "meta"
}