{
  "id": 499757,
  "title": "PyTorch Lightning mAP metric - Trick to reduce RAM usage",
  "url": "/competitions/leash-BELKA/discussion/499757",
  "author_name": "Luigi Stf",
  "post_date": "2024-05-02T22:18:18.476000",
  "votes": 17,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi fellows!</p>\n<p>I was playing with the awesome notebook <a href=\"https://www.kaggle.com/code/tetsuya3510/leash-bio-chemberta-baseline\" target=\"_blank\">leash-bio-chemberta-baseline</a> from <a href=\"https://www.kaggle.com/tetsuya3510\" target=\"_blank\">@tetsuya3510</a> and I stumbled on NaN issues when scaling up the dataset size (at fixed batch size).</p>\n<p>And I found the reason ! As some of you may encounter too, I share it here to avoid extra frustration 😁</p>\n<h3>Context</h3>\n<p>The PytorchLightning implementation of the mAP is <a href=\"https://lightning.ai/docs/torchmetrics/stable/classification/average_precision.html#\" target=\"_blank\">given here</a>.<br>\nIn the <a href=\"https://www.kaggle.com/tetsuya3510\" target=\"_blank\">@tetsuya3510</a> 's notebook, the mAP metric is thus created as follows :</p>\n<pre><code> (L.LightningModule):\n     ():\n        ().__init__()\n        self.model = LMModel(model_name)\n        self. = AveragePrecision(task=)\n</code></pre>\n<p>When using \"small datasets\" (100k rows), it worked perfectly. But when I scaled up to 1M rows, the validation mAP had weird values (NaN, -inf, or values not in 0/1). And surprisingly, when using Kaggle notebooks, everything went well = <strong>so it was an issue with my machine</strong></p>\n<h3>The issue</h3>\n<p>I then used debug mode and watched how the attributes of the metric evolved over time. And god, I found some interesting stuff !</p>\n<p>To compute the mAP, Lightning saves ALL THE PREDICTIONS+LABELS OF THE WHOLE EPOCH … to accurately calculate the mAP. </p>\n<p>Because of my limited resources, my memory couldn't handle so much tensors, so my machine failed.</p>\n<h3>The solution 🥳</h3>\n<p>The solution was in the official doc : we can approximate it, using bins, without saving everything !</p>\n<blockquote>\n  <p>The implementation both supports calculating the metric in a non-binned but accurate version and a binned version that is less accurate but more memory efficient. Setting the thresholds argument to None will activate the non-binned version that uses memory of size $O(n_samples)$ whereas setting the thresholds argument to either an integer, list or a 1d tensor will use a binned version that uses memory of size $O(n_thresholds)$ (constant memory).</p>\n</blockquote>\n<p>So provide the mAP object some thresholds, and your memory usage will shrink drastically ! Note that if you want to use the exact calculation, changing the type of the tensors to lower precisions is also an intermediate option.</p>\n<pre><code> (L.LightningModule):\n     ():\n        ().__init__()\n        self.model = LMModel(model_name)\n        self. = AveragePrecision(task=, thresholds=) \n</code></pre>\n<hr>\n<p>Voila. I hope it can help ! I almost lost 2 days on this issue, looking for a loophole in my code …</p>\n<p>I wish you happy kaggling </p>",
  "messages": [
    {
      "id": 2789901,
      "postDate": "2024-05-02T22:18:18.477Z",
      "content": "<p>Hi fellows!</p>\n<p>I was playing with the awesome notebook <a href=\"https://www.kaggle.com/code/tetsuya3510/leash-bio-chemberta-baseline\" target=\"_blank\">leash-bio-chemberta-baseline</a> from <a href=\"https://www.kaggle.com/tetsuya3510\" target=\"_blank\">@tetsuya3510</a> and I stumbled on NaN issues when scaling up the dataset size (at fixed batch size).</p>\n<p>And I found the reason ! As some of you may encounter too, I share it here to avoid extra frustration 😁</p>\n<h3>Context</h3>\n<p>The PytorchLightning implementation of the mAP is <a href=\"https://lightning.ai/docs/torchmetrics/stable/classification/average_precision.html#\" target=\"_blank\">given here</a>.<br>\nIn the <a href=\"https://www.kaggle.com/tetsuya3510\" target=\"_blank\">@tetsuya3510</a> 's notebook, the mAP metric is thus created as follows :</p>\n<pre><code> (L.LightningModule):\n     ():\n        ().__init__()\n        self.model = LMModel(model_name)\n        self. = AveragePrecision(task=)\n</code></pre>\n<p>When using \"small datasets\" (100k rows), it worked perfectly. But when I scaled up to 1M rows, the validation mAP had weird values (NaN, -inf, or values not in 0/1). And surprisingly, when using Kaggle notebooks, everything went well = <strong>so it was an issue with my machine</strong></p>\n<h3>The issue</h3>\n<p>I then used debug mode and watched how the attributes of the metric evolved over time. And god, I found some interesting stuff !</p>\n<p>To compute the mAP, Lightning saves ALL THE PREDICTIONS+LABELS OF THE WHOLE EPOCH … to accurately calculate the mAP. </p>\n<p>Because of my limited resources, my memory couldn't handle so much tensors, so my machine failed.</p>\n<h3>The solution 🥳</h3>\n<p>The solution was in the official doc : we can approximate it, using bins, without saving everything !</p>\n<blockquote>\n  <p>The implementation both supports calculating the metric in a non-binned but accurate version and a binned version that is less accurate but more memory efficient. Setting the thresholds argument to None will activate the non-binned version that uses memory of size $O(n_samples)$ whereas setting the thresholds argument to either an integer, list or a 1d tensor will use a binned version that uses memory of size $O(n_thresholds)$ (constant memory).</p>\n</blockquote>\n<p>So provide the mAP object some thresholds, and your memory usage will shrink drastically ! Note that if you want to use the exact calculation, changing the type of the tensors to lower precisions is also an intermediate option.</p>\n<pre><code> (L.LightningModule):\n     ():\n        ().__init__()\n        self.model = LMModel(model_name)\n        self. = AveragePrecision(task=, thresholds=) \n</code></pre>\n<hr>\n<p>Voila. I hope it can help ! I almost lost 2 days on this issue, looking for a loophole in my code …</p>\n<p>I wish you happy kaggling </p>",
      "rawMarkdown": "Hi fellows!\n\nI was playing with the awesome notebook [leash-bio-chemberta-baseline](https://www.kaggle.com/code/tetsuya3510/leash-bio-chemberta-baseline) from @tetsuya3510 and I stumbled on NaN issues when scaling up the dataset size (at fixed batch size).\n\nAnd I found the reason ! As some of you may encounter too, I share it here to avoid extra frustration 😁\n\n### Context\n\nThe PytorchLightning implementation of the mAP is [given here](https://lightning.ai/docs/torchmetrics/stable/classification/average_precision.html#).\nIn the @tetsuya3510 's notebook, the mAP metric is thus created as follows :\n\n```python\nclass LBModelModule(L.LightningModule):\n    def __init__(self, model_name):\n        super().__init__()\n        self.model = LMModel(model_name)\n        self.map = AveragePrecision(task=\"binary\")\n```\n\nWhen using \"small datasets\" (100k rows), it worked perfectly. But when I scaled up to 1M rows, the validation mAP had weird values (NaN, -inf, or values not in 0/1). And surprisingly, when using Kaggle notebooks, everything went well = **so it was an issue with my machine**\n\n### The issue\n\nI then used debug mode and watched how the attributes of the metric evolved over time. And god, I found some interesting stuff !\n\nTo compute the mAP, Lightning saves ALL THE PREDICTIONS+LABELS OF THE WHOLE EPOCH ... to accurately calculate the mAP. \n\nBecause of my limited resources, my memory couldn't handle so much tensors, so my machine failed.\n\n### The solution 🥳\n\nThe solution was in the official doc : we can approximate it, using bins, without saving everything !\n\n>The implementation both supports calculating the metric in a non-binned but accurate version and a binned version that is less accurate but more memory efficient. Setting the thresholds argument to None will activate the non-binned version that uses memory of size $O(n_samples)$ whereas setting the thresholds argument to either an integer, list or a 1d tensor will use a binned version that uses memory of size $O(n_thresholds)$ (constant memory).\n\nSo provide the mAP object some thresholds, and your memory usage will shrink drastically ! Note that if you want to use the exact calculation, changing the type of the tensors to lower precisions is also an intermediate option.\n\n```python\nclass LBModelModule(L.LightningModule):\n    def __init__(self, model_name):\n        super().__init__()\n        self.model = LMModel(model_name)\n        self.map = AveragePrecision(task=\"binary\", thresholds=100) # will use 100 bins\n```\n\n---\n\nVoila. I hope it can help ! I almost lost 2 days on this issue, looking for a loophole in my code ...\n\nI wish you happy kaggling ",
      "votes": 17
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2789901": "Hi fellows!\n\nI was playing with the awesome notebook [leash-bio-chemberta-baseline](https://www.kaggle.com/code/tetsuya3510/leash-bio-chemberta-baseline) from @tetsuya3510 and I stumbled on NaN issues when scaling up the dataset size (at fixed batch size).\n\nAnd I found the reason ! As some of you may encounter too, I share it here to avoid extra frustration 😁\n\n### Context\n\nThe PytorchLightning implementation of the mAP is [given here](https://lightning.ai/docs/torchmetrics/stable/classification/average_precision.html#).\nIn the @tetsuya3510 's notebook, the mAP metric is thus created as follows :\n\n```python\nclass LBModelModule(L.LightningModule):\n    def __init__(self, model_name):\n        super().__init__()\n        self.model = LMModel(model_name)\n        self.map = AveragePrecision(task=\"binary\")\n```\n\nWhen using \"small datasets\" (100k rows), it worked perfectly. But when I scaled up to 1M rows, the validation mAP had weird values (NaN, -inf, or values not in 0/1). And surprisingly, when using Kaggle notebooks, everything went well = **so it was an issue with my machine**\n\n### The issue\n\nI then used debug mode and watched how the attributes of the metric evolved over time. And god, I found some interesting stuff !\n\nTo compute the mAP, Lightning saves ALL THE PREDICTIONS+LABELS OF THE WHOLE EPOCH ... to accurately calculate the mAP. \n\nBecause of my limited resources, my memory couldn't handle so much tensors, so my machine failed.\n\n### The solution 🥳\n\nThe solution was in the official doc : we can approximate it, using bins, without saving everything !\n\n>The implementation both supports calculating the metric in a non-binned but accurate version and a binned version that is less accurate but more memory efficient. Setting the thresholds argument to None will activate the non-binned version that uses memory of size $O(n_samples)$ whereas setting the thresholds argument to either an integer, list or a 1d tensor will use a binned version that uses memory of size $O(n_thresholds)$ (constant memory).\n\nSo provide the mAP object some thresholds, and your memory usage will shrink drastically ! Note that if you want to use the exact calculation, changing the type of the tensors to lower precisions is also an intermediate option.\n\n```python\nclass LBModelModule(L.LightningModule):\n    def __init__(self, model_name):\n        super().__init__()\n        self.model = LMModel(model_name)\n        self.map = AveragePrecision(task=\"binary\", thresholds=100) # will use 100 bins\n```\n\n---\n\nVoila. I hope it can help ! I almost lost 2 days on this issue, looking for a loophole in my code ...\n\nI wish you happy kaggling "
  }
}