{
  "id": 454546,
  "title": "Intuition behind clipped/constrained predictions not training. ",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/454546",
  "author_name": "",
  "post_date": "2023-11-10T17:09:33.294036200Z",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>So the loss function I am using looks something like below (minus some masking shenanigans) and my model trains just fine: </p>\n<pre><code>def loss:\n    pred = pred\n    target  = batch.clip(,)\n\n    loss = l1    \n    return loss\n</code></pre>\n<p>I though it might help to constrain my predictions to <code>(0, 1)</code> like below: </p>\n<pre><code>def loss:\n    pred = pred.clip(,)\n    target  = batch.clip(,)\n\n    loss = l1    \n    return loss\n</code></pre>\n<p>But my model fails to train properly in this case. <br>\nAlso, if I instead of clipping my predictions limit them to <code>(0, 1)</code> with a sigmoid function I get the same thing: a complete failure to learn. </p>\n<p>I find this a little unintuitive.<br>\nAnyone have an idea of why this might be happening?   </p>",
  "messages": [
    {
      "id": "2520229",
      "postDate": "11/10/2023 17:09:33",
      "content": "<p>So the loss function I am using looks something like below (minus some masking shenanigans) and my model trains just fine: </p>\n<pre><code>def loss:\n    pred = pred\n    target  = batch.clip(,)\n\n    loss = l1    \n    return loss\n</code></pre>\n<p>I though it might help to constrain my predictions to <code>(0, 1)</code> like below: </p>\n<pre><code>def loss:\n    pred = pred.clip(,)\n    target  = batch.clip(,)\n\n    loss = l1    \n    return loss\n</code></pre>\n<p>But my model fails to train properly in this case. <br>\nAlso, if I instead of clipping my predictions limit them to <code>(0, 1)</code> with a sigmoid function I get the same thing: a complete failure to learn. </p>\n<p>I find this a little unintuitive.<br>\nAnyone have an idea of why this might be happening?   </p>",
      "rawMarkdown": "So the loss function I am using looks something like below (minus some masking shenanigans) and my model trains just fine: \n```\ndef loss_fun(pred,batch):\n    pred = pred\n    target  = batch['targets'].clip(0,1)\n\n    loss = F.l1_loss(pred, target, reduction='mean')    \n    return loss\n```\nI though it might help to constrain my predictions to `(0, 1)` like below: \n\n```\ndef loss_fun(pred,batch):\n    pred = pred.clip(0,1)\n    target  = batch['targets'].clip(0,1)\n\n    loss = F.l1_loss(pred, target, reduction='mean')    \n    return loss\n```\n\nBut my model fails to train properly in this case. \nAlso, if I instead of clipping my predictions limit them to `(0, 1)` with a sigmoid function I get the same thing: a complete failure to learn. \n\nI find this a little unintuitive.\nAnyone have an idea of why this might be happening?",
      "votes": null
    },
    {
      "id": "2520601",
      "postDate": "11/11/2023 01:33:25",
      "content": "<p>My guess is that clipping might not be differentiable and would thus not allow to compute the gradient of your loss for the backpropagation. <br>\nSigmoid is better in that regard, but I think it's better to not constrain the model too much if you have enough data, and rather let it learn the range (0, 1) from data.<br>\nYou could clip the prediction before the metric though, to improve it a bit</p>",
      "rawMarkdown": "My guess is that clipping might not be differentiable and would thus not allow to compute the gradient of your loss for the backpropagation. \nSigmoid is better in that regard, but I think it's better to not constrain the model too much if you have enough data, and rather let it learn the range (0, 1) from data.\nYou could clip the prediction before the metric though, to improve it a bit",
      "votes": null
    },
    {
      "id": "2521489",
      "postDate": "11/11/2023 18:45:24",
      "content": "<p>Case 1. </p>\n<p>your model is learning using the first loss,<br>\npred =&gt; target <br>\n1.5 =&gt; 1.0, loss = 0.5<br>\n1.9 =&gt; 1.0, loss = 0.9<br>\n2.9 =&gt; 1.0, loss = 1.9</p>\n<p>Case 2.</p>\n<p>your model is learning using the second loss,<br>\npred =&gt; target <br>\n1.5 =&gt; 1.0, loss = 0.0<br>\n1.9 =&gt; 1.0, loss = 0.0<br>\n2.9 =&gt; 1.0, loss = 0.0</p>\n<p>Which case has more gradation?</p>\n<p>It is quite intuitive that the model with the second loss learns less.</p>",
      "rawMarkdown": "Case 1. \n\nyour model is learning using the first loss,\npred => target \n1.5 => 1.0, loss = 0.5\n1.9 => 1.0, loss = 0.9\n2.9 => 1.0, loss = 1.9\n\nCase 2.\n\nyour model is learning using the second loss,\npred => target \n1.5 => 1.0, loss = 0.0\n1.9 => 1.0, loss = 0.0\n2.9 => 1.0, loss = 0.0\n\nWhich case has more gradation?\n\nIt is quite intuitive that the model with the second loss learns less.",
      "votes": null
    },
    {
      "id": "2521619",
      "postDate": "11/11/2023 22:21:11",
      "content": "<p>When we design Generative models, if the generator outperforms the discriminator, it generates only limited outputs. Google \"mode collapse\". In that case, we intentionally crop the prediction of it, so that it generates more various outputs.</p>",
      "rawMarkdown": "When we design Generative models, if the generator outperforms the discriminator, it generates only limited outputs. Google \"mode collapse\". In that case, we intentionally crop the prediction of it, so that it generates more various outputs.",
      "votes": null
    },
    {
      "id": "2522007",
      "postDate": "11/12/2023 09:31:18",
      "content": "<p>Great, yeah I get that. <br>\nI think this might be part of it, although in this case (unless all values are 1) I would expect worse training, not no training. </p>\n<p>I see the same phenomenon when not clipping the targets either. </p>\n<p>Like you mentioned in your other comment, I suspected mode collapse as well, as if enough targets are 1/0 it might just default to predicting those and get stuck in a local minimum. </p>\n<p>It does seems that models here are prone to just predicting the means of each class</p>",
      "rawMarkdown": "Great, yeah I get that. \nI think this might be part of it, although in this case (unless all values are 1) I would expect worse training, not no training. \n\nI see the same phenomenon when not clipping the targets either. \n\nLike you mentioned in your other comment, I suspected mode collapse as well, as if enough targets are 1/0 it might just default to predicting those and get stuck in a local minimum. \n\nIt does seems that models here are prone to just predicting the means of each class",
      "votes": null
    },
    {
      "id": "2522282",
      "postDate": "11/12/2023 15:11:08",
      "content": "<p>Hey!</p>\n<p>Yeah, I definitely clip before the metric (and prediction!)</p>",
      "rawMarkdown": "Hey!\n\nYeah, I definitely clip before the metric (and prediction!)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2520601,
      "author_name": "albericlajarte",
      "author_url": "",
      "post_date": "11/11/2023 01:33:25",
      "content": "<p>My guess is that clipping might not be differentiable and would thus not allow to compute the gradient of your loss for the backpropagation. <br>\nSigmoid is better in that regard, but I think it's better to not constrain the model too much if you have enough data, and rather let it learn the range (0, 1) from data.<br>\nYou could clip the prediction before the metric though, to improve it a bit</p>",
      "votes": null,
      "replies": [
        {
          "id": 2522282,
          "author_name": "fnands",
          "author_url": "",
          "post_date": "11/12/2023 15:11:08",
          "content": "<p>Hey!</p>\n<p>Yeah, I definitely clip before the metric (and prediction!)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2521489,
      "author_name": "",
      "author_url": "",
      "post_date": "11/11/2023 18:45:24",
      "content": "<p>Case 1. </p>\n<p>your model is learning using the first loss,<br>\npred =&gt; target <br>\n1.5 =&gt; 1.0, loss = 0.5<br>\n1.9 =&gt; 1.0, loss = 0.9<br>\n2.9 =&gt; 1.0, loss = 1.9</p>\n<p>Case 2.</p>\n<p>your model is learning using the second loss,<br>\npred =&gt; target <br>\n1.5 =&gt; 1.0, loss = 0.0<br>\n1.9 =&gt; 1.0, loss = 0.0<br>\n2.9 =&gt; 1.0, loss = 0.0</p>\n<p>Which case has more gradation?</p>\n<p>It is quite intuitive that the model with the second loss learns less.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2522007,
          "author_name": "fnands",
          "author_url": "",
          "post_date": "11/12/2023 09:31:18",
          "content": "<p>Great, yeah I get that. <br>\nI think this might be part of it, although in this case (unless all values are 1) I would expect worse training, not no training. </p>\n<p>I see the same phenomenon when not clipping the targets either. </p>\n<p>Like you mentioned in your other comment, I suspected mode collapse as well, as if enough targets are 1/0 it might just default to predicting those and get stuck in a local minimum. </p>\n<p>It does seems that models here are prone to just predicting the means of each class</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2521619,
      "author_name": "",
      "author_url": "",
      "post_date": "11/11/2023 22:21:11",
      "content": "<p>When we design Generative models, if the generator outperforms the discriminator, it generates only limited outputs. Google \"mode collapse\". In that case, we intentionally crop the prediction of it, so that it generates more various outputs.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2520229": "So the loss function I am using looks something like below (minus some masking shenanigans) and my model trains just fine: \n```\ndef loss_fun(pred,batch):\n    pred = pred\n    target  = batch['targets'].clip(0,1)\n\n    loss = F.l1_loss(pred, target, reduction='mean')    \n    return loss\n```\nI though it might help to constrain my predictions to `(0, 1)` like below: \n\n```\ndef loss_fun(pred,batch):\n    pred = pred.clip(0,1)\n    target  = batch['targets'].clip(0,1)\n\n    loss = F.l1_loss(pred, target, reduction='mean')    \n    return loss\n```\n\nBut my model fails to train properly in this case. \nAlso, if I instead of clipping my predictions limit them to `(0, 1)` with a sigmoid function I get the same thing: a complete failure to learn. \n\nI find this a little unintuitive.\nAnyone have an idea of why this might be happening?",
    "2520601": "My guess is that clipping might not be differentiable and would thus not allow to compute the gradient of your loss for the backpropagation. \nSigmoid is better in that regard, but I think it's better to not constrain the model too much if you have enough data, and rather let it learn the range (0, 1) from data.\nYou could clip the prediction before the metric though, to improve it a bit",
    "2521489": "Case 1. \n\nyour model is learning using the first loss,\npred => target \n1.5 => 1.0, loss = 0.5\n1.9 => 1.0, loss = 0.9\n2.9 => 1.0, loss = 1.9\n\nCase 2.\n\nyour model is learning using the second loss,\npred => target \n1.5 => 1.0, loss = 0.0\n1.9 => 1.0, loss = 0.0\n2.9 => 1.0, loss = 0.0\n\nWhich case has more gradation?\n\nIt is quite intuitive that the model with the second loss learns less.",
    "2521619": "When we design Generative models, if the generator outperforms the discriminator, it generates only limited outputs. Google \"mode collapse\". In that case, we intentionally crop the prediction of it, so that it generates more various outputs.",
    "2522007": "Great, yeah I get that. \nI think this might be part of it, although in this case (unless all values are 1) I would expect worse training, not no training. \n\nI see the same phenomenon when not clipping the targets either. \n\nLike you mentioned in your other comment, I suspected mode collapse as well, as if enough targets are 1/0 it might just default to predicting those and get stuck in a local minimum. \n\nIt does seems that models here are prone to just predicting the means of each class",
    "2522282": "Hey!\n\nYeah, I definitely clip before the metric (and prediction!)"
  },
  "source": "meta"
}