{
  "id": 442490,
  "title": "Funky effects of batch size on the metric MRRMSE (You'd want to know if you are using NN)",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/442490",
  "author_name": "",
  "post_date": "2023-09-22T20:43:42.305481600Z",
  "votes": 7,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I was trying to directly minimize the MRRMSE using neural nets and found some really annoying characteristic of the MRRMSE (possibly because large values are sparse in the prediction targets?)</p>\n<p>If you have a small batch size (sometimes you have to because the neural net is large), you calcualte the MRRMSE by batches and the do the average. You will find that the MRRMSE would seem artificially lower when the batch size is small.</p>\n<p>This could be very confusing if you are not aware of this. (confused me half a day wondering why two models of similar capacity have vastly different validation loss, it turns out just due to different batch size in validation runs).</p>\n<p>It could be annoying too, because you need to calculate the loss for each mini batch, but then the loss is a bad proxy when the batch size is small.</p>\n<p>Here's what the MRRMSE would be, given the same prediction (all 0s) sampled in different batch sizes. (26 DE predictions for each gene in the chart, e.g. batch size of 100 genes is equvilant to 2600 DE predictions if your model predicts at the compound-cell-gene level)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1716193%2F42e0e938046444edacc49fdca36d7d23%2FScreenshot%202023-09-22%20163845.png?generation=1695415165222004&amp;alt=media\" alt=\"\"></p>\n<p>Take away, just use larger batch size if you are running neural nets.</p>",
  "messages": [
    {
      "id": "2451846",
      "postDate": "09/22/2023 20:43:42",
      "content": "<p>I was trying to directly minimize the MRRMSE using neural nets and found some really annoying characteristic of the MRRMSE (possibly because large values are sparse in the prediction targets?)</p>\n<p>If you have a small batch size (sometimes you have to because the neural net is large), you calcualte the MRRMSE by batches and the do the average. You will find that the MRRMSE would seem artificially lower when the batch size is small.</p>\n<p>This could be very confusing if you are not aware of this. (confused me half a day wondering why two models of similar capacity have vastly different validation loss, it turns out just due to different batch size in validation runs).</p>\n<p>It could be annoying too, because you need to calculate the loss for each mini batch, but then the loss is a bad proxy when the batch size is small.</p>\n<p>Here's what the MRRMSE would be, given the same prediction (all 0s) sampled in different batch sizes. (26 DE predictions for each gene in the chart, e.g. batch size of 100 genes is equvilant to 2600 DE predictions if your model predicts at the compound-cell-gene level)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1716193%2F42e0e938046444edacc49fdca36d7d23%2FScreenshot%202023-09-22%20163845.png?generation=1695415165222004&amp;alt=media\" alt=\"\"></p>\n<p>Take away, just use larger batch size if you are running neural nets.</p>",
      "rawMarkdown": "I was trying to directly minimize the MRRMSE using neural nets and found some really annoying characteristic of the MRRMSE (possibly because large values are sparse in the prediction targets?)\n\nIf you have a small batch size (sometimes you have to because the neural net is large), you calcualte the MRRMSE by batches and the do the average. You will find that the MRRMSE would seem artificially lower when the batch size is small.\n\nThis could be very confusing if you are not aware of this. (confused me half a day wondering why two models of similar capacity have vastly different validation loss, it turns out just due to different batch size in validation runs).\n\nIt could be annoying too, because you need to calculate the loss for each mini batch, but then the loss is a bad proxy when the batch size is small.\n\nHere's what the MRRMSE would be, given the same prediction (all 0s) sampled in different batch sizes. (26 DE predictions for each gene in the chart, e.g. batch size of 100 genes is equvilant to 2600 DE predictions if your model predicts at the compound-cell-gene level)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1716193%2F42e0e938046444edacc49fdca36d7d23%2FScreenshot%202023-09-22%20163845.png?generation=1695415165222004&alt=media)\n\n\nTake away, just use larger batch size if you are running neural nets.",
      "votes": null
    },
    {
      "id": "2454209",
      "postDate": "09/24/2023 16:35:32",
      "content": "<p>If you want to compare two different models, you can just use MSE or RMSE, which are invariant to batch size,<br>\nIf MSE for model1 is lower than MSE for model2, MRRMSE for model1 will be lower than MRRMSE for model2, and MSE is not dependent on batch size. <br>\nI tried to verify this experimentally, but it turns out that if the RMSE losses are very close (within ~1e-6) floating point error breaks this relationship down.</p>\n<pre><code>import torch\n\nrows = \ncols = \nactual = torch.rand(rows,cols,dtype=torch.float64)\n\ndef vectorized_mrrmse(,actual):\n     = torch.square( - actual)\n    col_sum = torch.(,axis=)/cols\n    row_sum = torch.(torch.(col_sum))/rows\n     row_sum\n\n\ndef naive_mrrmse(,actual):\n    t = \n     i  (rows):\n        s = \n         j  (cols):\n            s += torch.square([i,j] - actual[i,j])\n        t += torch.(s/cols)\n     t/rows\n\n\ndef batched(,actual,bsz,lsfn):\n    ac_batch = torch.(actual,bsz)\n    pred_batch = torch.(,bsz)\n    assert len(ac_batch) == len(pred_batch)\n    vals = []\n     b  (len(ac_batch)):\n        # Batches can be different sizes\n        bsz = pred_batch[b].shape[]\n        vals.(bsz*lsfn(pred_batch[b],ac_batch[b]))\n    vals = torch.stack(vals)\n     torch.(vals) / .shape[]\n\nmsels = torch.nn.MSELoss()\ndef rmse(,actual):\n     torch.(msels(,actual))\n\n i  ():\n    pred1 = torch.rand(rows,cols,dtype=torch.float64)\n    pred2 = torch.rand(rows,cols,dtype=torch.float64)\n    (i)\n     rmse(pred1,actual) &lt; rmse(pred2,actual):\n          vectorized_mrrmse(pred1,actual) &lt; vectorized_mrrmse(pred2,actual):\n             = vectorized_mrrmse(pred1,actual) - vectorized_mrrmse(pred2,actual)\n            ()\n            exit()\n    :\n          vectorized_mrrmse(pred1,actual) &gt; vectorized_mrrmse(pred2,actual):\n             = vectorized_mrrmse(pred1,actual) - vectorized_mrrmse(pred2,actual)\n            ()\n            exit()\n</code></pre>\n<p>Code is pretty sloppy.</p>",
      "rawMarkdown": "If you want to compare two different models, you can just use MSE or RMSE, which are invariant to batch size,\nIf MSE for model1 is lower than MSE for model2, MRRMSE for model1 will be lower than MRRMSE for model2, and MSE is not dependent on batch size. \nI tried to verify this experimentally, but it turns out that if the RMSE losses are very close (within ~1e-6) floating point error breaks this relationship down.\n\n```\nimport torch\n\nrows = 110\ncols = 3000\nactual = torch.rand(rows,cols,dtype=torch.float64)\n\ndef vectorized_mrrmse(pred,actual):\n    diff = torch.square(pred - actual)\n    col_sum = torch.sum(diff,axis=1)/cols\n    row_sum = torch.sum(torch.sqrt(col_sum))/rows\n    return row_sum\n\n\ndef naive_mrrmse(pred,actual):\n    t = 0\n    for i in range(rows):\n        s = 0\n        for j in range(cols):\n            s += torch.square(pred[i,j] - actual[i,j])\n        t += torch.sqrt(s/cols)\n    return t/rows\n\n\ndef batched(pred,actual,bsz,lsfn):\n    ac_batch = torch.split(actual,bsz)\n    pred_batch = torch.split(pred,bsz)\n    assert len(ac_batch) == len(pred_batch)\n    vals = []\n    for b in range(len(ac_batch)):\n        # Batches can be different sizes\n        bsz = pred_batch[b].shape[0]\n        vals.append(bsz*lsfn(pred_batch[b],ac_batch[b]))\n    vals = torch.stack(vals)\n    return torch.sum(vals) / pred.shape[0]\n\nmsels = torch.nn.MSELoss()\ndef rmse(pred,actual):\n    return torch.sqrt(msels(pred,actual))\n\nfor i in range(10000):\n    pred1 = torch.rand(rows,cols,dtype=torch.float64)\n    pred2 = torch.rand(rows,cols,dtype=torch.float64)\n    print(i)\n    if rmse(pred1,actual) < rmse(pred2,actual):\n        if not vectorized_mrrmse(pred1,actual) < vectorized_mrrmse(pred2,actual):\n            error = vectorized_mrrmse(pred1,actual) - vectorized_mrrmse(pred2,actual)\n            print(error)\n            exit()\n    else:\n        if not vectorized_mrrmse(pred1,actual) > vectorized_mrrmse(pred2,actual):\n            error = vectorized_mrrmse(pred1,actual) - vectorized_mrrmse(pred2,actual)\n            print(error)\n            exit()\n```\nCode is pretty sloppy.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2454209,
      "author_name": "laurasisson",
      "author_url": "",
      "post_date": "09/24/2023 16:35:32",
      "content": "<p>If you want to compare two different models, you can just use MSE or RMSE, which are invariant to batch size,<br>\nIf MSE for model1 is lower than MSE for model2, MRRMSE for model1 will be lower than MRRMSE for model2, and MSE is not dependent on batch size. <br>\nI tried to verify this experimentally, but it turns out that if the RMSE losses are very close (within ~1e-6) floating point error breaks this relationship down.</p>\n<pre><code>import torch\n\nrows = \ncols = \nactual = torch.rand(rows,cols,dtype=torch.float64)\n\ndef vectorized_mrrmse(,actual):\n     = torch.square( - actual)\n    col_sum = torch.(,axis=)/cols\n    row_sum = torch.(torch.(col_sum))/rows\n     row_sum\n\n\ndef naive_mrrmse(,actual):\n    t = \n     i  (rows):\n        s = \n         j  (cols):\n            s += torch.square([i,j] - actual[i,j])\n        t += torch.(s/cols)\n     t/rows\n\n\ndef batched(,actual,bsz,lsfn):\n    ac_batch = torch.(actual,bsz)\n    pred_batch = torch.(,bsz)\n    assert len(ac_batch) == len(pred_batch)\n    vals = []\n     b  (len(ac_batch)):\n        # Batches can be different sizes\n        bsz = pred_batch[b].shape[]\n        vals.(bsz*lsfn(pred_batch[b],ac_batch[b]))\n    vals = torch.stack(vals)\n     torch.(vals) / .shape[]\n\nmsels = torch.nn.MSELoss()\ndef rmse(,actual):\n     torch.(msels(,actual))\n\n i  ():\n    pred1 = torch.rand(rows,cols,dtype=torch.float64)\n    pred2 = torch.rand(rows,cols,dtype=torch.float64)\n    (i)\n     rmse(pred1,actual) &lt; rmse(pred2,actual):\n          vectorized_mrrmse(pred1,actual) &lt; vectorized_mrrmse(pred2,actual):\n             = vectorized_mrrmse(pred1,actual) - vectorized_mrrmse(pred2,actual)\n            ()\n            exit()\n    :\n          vectorized_mrrmse(pred1,actual) &gt; vectorized_mrrmse(pred2,actual):\n             = vectorized_mrrmse(pred1,actual) - vectorized_mrrmse(pred2,actual)\n            ()\n            exit()\n</code></pre>\n<p>Code is pretty sloppy.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2451846": "I was trying to directly minimize the MRRMSE using neural nets and found some really annoying characteristic of the MRRMSE (possibly because large values are sparse in the prediction targets?)\n\nIf you have a small batch size (sometimes you have to because the neural net is large), you calcualte the MRRMSE by batches and the do the average. You will find that the MRRMSE would seem artificially lower when the batch size is small.\n\nThis could be very confusing if you are not aware of this. (confused me half a day wondering why two models of similar capacity have vastly different validation loss, it turns out just due to different batch size in validation runs).\n\nIt could be annoying too, because you need to calculate the loss for each mini batch, but then the loss is a bad proxy when the batch size is small.\n\nHere's what the MRRMSE would be, given the same prediction (all 0s) sampled in different batch sizes. (26 DE predictions for each gene in the chart, e.g. batch size of 100 genes is equvilant to 2600 DE predictions if your model predicts at the compound-cell-gene level)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1716193%2F42e0e938046444edacc49fdca36d7d23%2FScreenshot%202023-09-22%20163845.png?generation=1695415165222004&alt=media)\n\n\nTake away, just use larger batch size if you are running neural nets.",
    "2454209": "If you want to compare two different models, you can just use MSE or RMSE, which are invariant to batch size,\nIf MSE for model1 is lower than MSE for model2, MRRMSE for model1 will be lower than MRRMSE for model2, and MSE is not dependent on batch size. \nI tried to verify this experimentally, but it turns out that if the RMSE losses are very close (within ~1e-6) floating point error breaks this relationship down.\n\n```\nimport torch\n\nrows = 110\ncols = 3000\nactual = torch.rand(rows,cols,dtype=torch.float64)\n\ndef vectorized_mrrmse(pred,actual):\n    diff = torch.square(pred - actual)\n    col_sum = torch.sum(diff,axis=1)/cols\n    row_sum = torch.sum(torch.sqrt(col_sum))/rows\n    return row_sum\n\n\ndef naive_mrrmse(pred,actual):\n    t = 0\n    for i in range(rows):\n        s = 0\n        for j in range(cols):\n            s += torch.square(pred[i,j] - actual[i,j])\n        t += torch.sqrt(s/cols)\n    return t/rows\n\n\ndef batched(pred,actual,bsz,lsfn):\n    ac_batch = torch.split(actual,bsz)\n    pred_batch = torch.split(pred,bsz)\n    assert len(ac_batch) == len(pred_batch)\n    vals = []\n    for b in range(len(ac_batch)):\n        # Batches can be different sizes\n        bsz = pred_batch[b].shape[0]\n        vals.append(bsz*lsfn(pred_batch[b],ac_batch[b]))\n    vals = torch.stack(vals)\n    return torch.sum(vals) / pred.shape[0]\n\nmsels = torch.nn.MSELoss()\ndef rmse(pred,actual):\n    return torch.sqrt(msels(pred,actual))\n\nfor i in range(10000):\n    pred1 = torch.rand(rows,cols,dtype=torch.float64)\n    pred2 = torch.rand(rows,cols,dtype=torch.float64)\n    print(i)\n    if rmse(pred1,actual) < rmse(pred2,actual):\n        if not vectorized_mrrmse(pred1,actual) < vectorized_mrrmse(pred2,actual):\n            error = vectorized_mrrmse(pred1,actual) - vectorized_mrrmse(pred2,actual)\n            print(error)\n            exit()\n    else:\n        if not vectorized_mrrmse(pred1,actual) > vectorized_mrrmse(pred2,actual):\n            error = vectorized_mrrmse(pred1,actual) - vectorized_mrrmse(pred2,actual)\n            print(error)\n            exit()\n```\nCode is pretty sloppy."
  },
  "source": "meta"
}