{
  "id": 180732,
  "title": "Separate FVC and Sigma prediction",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/180732",
  "author_name": "",
  "post_date": "2020-09-06T09:31:33.901405300Z",
  "votes": 7,
  "comment_count": 4,
  "views": 0,
  "content": "<p><strong>Laplace Loss Split</strong></p>\n<p>Why we want to split the competition into separate FVC and Sigma prediction? There are some reasons:</p>\n<ul>\n<li>For using boost algorithms with 1D model output (XGBoost, LightGBM, CatBoost).</li>\n<li>[More popular on my side] Absence of true Sigma labels.</li>\n</ul>\n<p>There is mathematical proof that splitting is the next:</p>\n<p>$$<br>\nLoss(y_{true}, y_{pred}, \\sigma_{pred}) = Loss_1(y_{true}, y_{pred}) + Loss_2(\\Delta_{true}, \\Delta_{pred}),<br>\n$$</p>\n<p>where </p>\n<p>$$<br>\n\\sigma = \\sqrt{2} \\cdot \\Delta_{pred}, \\<br>\n   \\Delta = \\Delta_{true}<br>\n$$</p>\n<p>So it works as follows: at first you train a model for FVC prediction, and than for these predictions compute optimal Sigma and train the second model. You can use it as a custom loss for your models or as a metric for understanding what model is the best after training. In my experiments first approach with boost algorithms haven't worked so I'm going to use the second option.</p>\n<p><strong>First Loss</strong></p>\n<p>$$<br>\nLoss_1(y_{true}, y_{pred}) = \\log(|y_{true} - y_{pred}|)<br>\n$$</p>\n<p>First loss equals <em>geometric mean</em> of the module of the errors, so behaviour is obvious:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2Fd73c2016accb4eff861c893b2eae8e15%2FLoss1.png?generation=1599383049802037&amp;alt=media\" alt=\"\"></p>\n<p><strong>Second Loss</strong></p>\n<p>$$<br>\nLoss_2(\\Delta_{true}, \\Delta_{pred}) = \\frac{\\Delta_{true}}{\\Delta_{pred}} + \\log{\\frac{\\Delta_{pred}}{\\Delta_{true}}}.<br>\n$$</p>\n<p>That loss is more sophisticated and has two cases that different in it's behaviour:</p>\n<p>$$ \\Delta_{pred} &lt; \\Delta_{true}, (x = \\frac{\\Delta_{true}}{\\Delta_{pred}} &gt; 1) \\rightarrow Loss_2(\\Delta_{true}, \\Delta_{pred}) = x - \\ln x$$<br>\n$$\\Delta_{pred} \\geq \\Delta_{true}, (x = \\frac{\\Delta_{pred}}{\\Delta_{true}} \\geq 1) \\rightarrow Loss_2(\\Delta_{true}, \\Delta_{pred}) = \\frac{1}{x} + \\ln x \\sim  \\ln x$$</p>\n<p>For instance, penalty for underrated Sigma is stronger, than for overrated one. It's clear to see on the pictures below (even with the next fact: more true Sigma, stronger penalty for underestimation):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2Fae6ccf733d088eedafbe4b4cbb9d82e2%2FLoss2.png?generation=1599383072353590&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "1000094",
      "postDate": "09/06/2020 09:31:33",
      "content": "<p><strong>Laplace Loss Split</strong></p>\n<p>Why we want to split the competition into separate FVC and Sigma prediction? There are some reasons:</p>\n<ul>\n<li>For using boost algorithms with 1D model output (XGBoost, LightGBM, CatBoost).</li>\n<li>[More popular on my side] Absence of true Sigma labels.</li>\n</ul>\n<p>There is mathematical proof that splitting is the next:</p>\n<p>$$<br>\nLoss(y_{true}, y_{pred}, \\sigma_{pred}) = Loss_1(y_{true}, y_{pred}) + Loss_2(\\Delta_{true}, \\Delta_{pred}),<br>\n$$</p>\n<p>where </p>\n<p>$$<br>\n\\sigma = \\sqrt{2} \\cdot \\Delta_{pred}, \\<br>\n   \\Delta = \\Delta_{true}<br>\n$$</p>\n<p>So it works as follows: at first you train a model for FVC prediction, and than for these predictions compute optimal Sigma and train the second model. You can use it as a custom loss for your models or as a metric for understanding what model is the best after training. In my experiments first approach with boost algorithms haven't worked so I'm going to use the second option.</p>\n<p><strong>First Loss</strong></p>\n<p>$$<br>\nLoss_1(y_{true}, y_{pred}) = \\log(|y_{true} - y_{pred}|)<br>\n$$</p>\n<p>First loss equals <em>geometric mean</em> of the module of the errors, so behaviour is obvious:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2Fd73c2016accb4eff861c893b2eae8e15%2FLoss1.png?generation=1599383049802037&amp;alt=media\" alt=\"\"></p>\n<p><strong>Second Loss</strong></p>\n<p>$$<br>\nLoss_2(\\Delta_{true}, \\Delta_{pred}) = \\frac{\\Delta_{true}}{\\Delta_{pred}} + \\log{\\frac{\\Delta_{pred}}{\\Delta_{true}}}.<br>\n$$</p>\n<p>That loss is more sophisticated and has two cases that different in it's behaviour:</p>\n<p>$$ \\Delta_{pred} &lt; \\Delta_{true}, (x = \\frac{\\Delta_{true}}{\\Delta_{pred}} &gt; 1) \\rightarrow Loss_2(\\Delta_{true}, \\Delta_{pred}) = x - \\ln x$$<br>\n$$\\Delta_{pred} \\geq \\Delta_{true}, (x = \\frac{\\Delta_{pred}}{\\Delta_{true}} \\geq 1) \\rightarrow Loss_2(\\Delta_{true}, \\Delta_{pred}) = \\frac{1}{x} + \\ln x \\sim  \\ln x$$</p>\n<p>For instance, penalty for underrated Sigma is stronger, than for overrated one. It's clear to see on the pictures below (even with the next fact: more true Sigma, stronger penalty for underestimation):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2Fae6ccf733d088eedafbe4b4cbb9d82e2%2FLoss2.png?generation=1599383072353590&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "**Laplace Loss Split**\n\nWhy we want to split the competition into separate FVC and Sigma prediction? There are some reasons:\n- For using boost algorithms with 1D model output (XGBoost, LightGBM, CatBoost).\n- [More popular on my side] Absence of true Sigma labels.\n\nThere is mathematical proof that splitting is the next:\n\n$$\nLoss(y_{true}, y_{pred}, \\sigma_{pred}) = Loss_1(y_{true}, y_{pred}) + Loss_2(\\Delta_{true}, \\Delta_{pred}),\n$$\n\nwhere \n\n$$\n\\sigma = \\sqrt{2} \\cdot \\Delta_{pred}, \\\\\n   \\Delta = \\Delta_{true}\n$$\n\nSo it works as follows: at first you train a model for FVC prediction, and than for these predictions compute optimal Sigma and train the second model. You can use it as a custom loss for your models or as a metric for understanding what model is the best after training. In my experiments first approach with boost algorithms haven't worked so I'm going to use the second option.\n\n**First Loss**\n\n$$\nLoss_1(y_{true}, y_{pred}) = \\log(|y_{true} - y_{pred}|)\n$$\n\nFirst loss equals *geometric mean* of the module of the errors, so behaviour is obvious:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2Fd73c2016accb4eff861c893b2eae8e15%2FLoss1.png?generation=1599383049802037&alt=media)\n\n**Second Loss**\n\n$$\nLoss_2(\\Delta_{true}, \\Delta_{pred}) = \\frac{\\Delta_{true}}{\\Delta_{pred}} + \\log{\\frac{\\Delta_{pred}}{\\Delta_{true}}}.\n$$\n\nThat loss is more sophisticated and has two cases that different in it's behaviour:\n\n$$ \\Delta_{pred} < \\Delta_{true}, (x = \\frac{\\Delta_{true}}{\\Delta_{pred}} > 1) \\rightarrow Loss_2(\\Delta_{true}, \\Delta_{pred}) = x - \\ln x$$\n$$\\Delta_{pred} \\geq \\Delta_{true}, (x = \\frac{\\Delta_{pred}}{\\Delta_{true}} \\geq 1) \\rightarrow Loss_2(\\Delta_{true}, \\Delta_{pred}) = \\frac{1}{x} + \\ln x \\sim  \\ln x$$\n\nFor instance, penalty for underrated Sigma is stronger, than for overrated one. It's clear to see on the pictures below (even with the next fact: more true Sigma, stronger penalty for underestimation):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2Fae6ccf733d088eedafbe4b4cbb9d82e2%2FLoss2.png?generation=1599383072353590&alt=media)",
      "votes": null
    },
    {
      "id": "1000108",
      "postDate": "09/06/2020 09:44:59",
      "content": "<p>what's  $\\Delta$  standards for? </p>\n<p>should the loss1 in your post be scaled by sigma? </p>",
      "rawMarkdown": "what's  $\\Delta$  standards for? \n\nshould the loss1 in your post be scaled by sigma?",
      "votes": null
    },
    {
      "id": "1000120",
      "postDate": "09/06/2020 09:58:29",
      "content": "<p>Delta is an abbreviation from the Laplace loss:<br>\n$$<br>\n\\Delta = \\min(|FVC_{true} - FVC_{predicted}, 1000)<br>\n$$</p>",
      "rawMarkdown": "Delta is an abbreviation from the Laplace loss:\n$$\n\\Delta = \\min(|FVC_{true} - FVC_{predicted}, 1000)\n$$",
      "votes": null
    },
    {
      "id": "1000196",
      "postDate": "09/06/2020 11:23:00",
      "content": "<pre><code>at first you train a model for FVC prediction, and than for these predictions compute optimal Sigma and train the second model.\n</code></pre>\n<p>This is what was done in the following notebook, correct? </p>\n<p><a href=\"https://www.kaggle.com/yasufuminakama/osic-lgb-baseline\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/osic-lgb-baseline</a></p>",
      "rawMarkdown": "```\nat first you train a model for FVC prediction, and than for these predictions compute optimal Sigma and train the second model.\n```\n\nThis is what was done in the following notebook, correct? \n\nhttps://www.kaggle.com/yasufuminakama/osic-lgb-baseline",
      "votes": null
    },
    {
      "id": "1000210",
      "postDate": "09/06/2020 11:37:13",
      "content": "<p>Great notebook, exactly the same idea. But this topic more about the losses: the notebook has used RMSE for each task (chances are a good practice solution indeed), but I suggest to use particular losses for each step with high correlation with Laplace Loss. For instance <code>RMSE=0.001</code> means that you get good model but you still don't know what Laplace Loss you get after all. On the other hand <code>Loss_1=4.321</code> means lower bound for Laplace Loss (drop some constant for better understanding).</p>",
      "rawMarkdown": "Great notebook, exactly the same idea. But this topic more about the losses: the notebook has used RMSE for each task (chances are a good practice solution indeed), but I suggest to use particular losses for each step with high correlation with Laplace Loss. For instance `RMSE=0.001` means that you get good model but you still don't know what Laplace Loss you get after all. On the other hand `Loss_1=4.321` means lower bound for Laplace Loss (drop some constant for better understanding).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1000108,
      "author_name": "yimacs",
      "author_url": "",
      "post_date": "09/06/2020 09:44:59",
      "content": "<p>what's  $\\Delta$  standards for? </p>\n<p>should the loss1 in your post be scaled by sigma? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1000120,
          "author_name": "koza4ukdmitrij",
          "author_url": "",
          "post_date": "09/06/2020 09:58:29",
          "content": "<p>Delta is an abbreviation from the Laplace loss:<br>\n$$<br>\n\\Delta = \\min(|FVC_{true} - FVC_{predicted}, 1000)<br>\n$$</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1000196,
      "author_name": "code1110",
      "author_url": "",
      "post_date": "09/06/2020 11:23:00",
      "content": "<pre><code>at first you train a model for FVC prediction, and than for these predictions compute optimal Sigma and train the second model.\n</code></pre>\n<p>This is what was done in the following notebook, correct? </p>\n<p><a href=\"https://www.kaggle.com/yasufuminakama/osic-lgb-baseline\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/osic-lgb-baseline</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1000210,
          "author_name": "koza4ukdmitrij",
          "author_url": "",
          "post_date": "09/06/2020 11:37:13",
          "content": "<p>Great notebook, exactly the same idea. But this topic more about the losses: the notebook has used RMSE for each task (chances are a good practice solution indeed), but I suggest to use particular losses for each step with high correlation with Laplace Loss. For instance <code>RMSE=0.001</code> means that you get good model but you still don't know what Laplace Loss you get after all. On the other hand <code>Loss_1=4.321</code> means lower bound for Laplace Loss (drop some constant for better understanding).</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1000094": "**Laplace Loss Split**\n\nWhy we want to split the competition into separate FVC and Sigma prediction? There are some reasons:\n- For using boost algorithms with 1D model output (XGBoost, LightGBM, CatBoost).\n- [More popular on my side] Absence of true Sigma labels.\n\nThere is mathematical proof that splitting is the next:\n\n$$\nLoss(y_{true}, y_{pred}, \\sigma_{pred}) = Loss_1(y_{true}, y_{pred}) + Loss_2(\\Delta_{true}, \\Delta_{pred}),\n$$\n\nwhere \n\n$$\n\\sigma = \\sqrt{2} \\cdot \\Delta_{pred}, \\\\\n   \\Delta = \\Delta_{true}\n$$\n\nSo it works as follows: at first you train a model for FVC prediction, and than for these predictions compute optimal Sigma and train the second model. You can use it as a custom loss for your models or as a metric for understanding what model is the best after training. In my experiments first approach with boost algorithms haven't worked so I'm going to use the second option.\n\n**First Loss**\n\n$$\nLoss_1(y_{true}, y_{pred}) = \\log(|y_{true} - y_{pred}|)\n$$\n\nFirst loss equals *geometric mean* of the module of the errors, so behaviour is obvious:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2Fd73c2016accb4eff861c893b2eae8e15%2FLoss1.png?generation=1599383049802037&alt=media)\n\n**Second Loss**\n\n$$\nLoss_2(\\Delta_{true}, \\Delta_{pred}) = \\frac{\\Delta_{true}}{\\Delta_{pred}} + \\log{\\frac{\\Delta_{pred}}{\\Delta_{true}}}.\n$$\n\nThat loss is more sophisticated and has two cases that different in it's behaviour:\n\n$$ \\Delta_{pred} < \\Delta_{true}, (x = \\frac{\\Delta_{true}}{\\Delta_{pred}} > 1) \\rightarrow Loss_2(\\Delta_{true}, \\Delta_{pred}) = x - \\ln x$$\n$$\\Delta_{pred} \\geq \\Delta_{true}, (x = \\frac{\\Delta_{pred}}{\\Delta_{true}} \\geq 1) \\rightarrow Loss_2(\\Delta_{true}, \\Delta_{pred}) = \\frac{1}{x} + \\ln x \\sim  \\ln x$$\n\nFor instance, penalty for underrated Sigma is stronger, than for overrated one. It's clear to see on the pictures below (even with the next fact: more true Sigma, stronger penalty for underestimation):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2Fae6ccf733d088eedafbe4b4cbb9d82e2%2FLoss2.png?generation=1599383072353590&alt=media)",
    "1000108": "what's  $\\Delta$  standards for? \n\nshould the loss1 in your post be scaled by sigma?",
    "1000120": "Delta is an abbreviation from the Laplace loss:\n$$\n\\Delta = \\min(|FVC_{true} - FVC_{predicted}, 1000)\n$$",
    "1000196": "```\nat first you train a model for FVC prediction, and than for these predictions compute optimal Sigma and train the second model.\n```\n\nThis is what was done in the following notebook, correct? \n\nhttps://www.kaggle.com/yasufuminakama/osic-lgb-baseline",
    "1000210": "Great notebook, exactly the same idea. But this topic more about the losses: the notebook has used RMSE for each task (chances are a good practice solution indeed), but I suggest to use particular losses for each step with high correlation with Laplace Loss. For instance `RMSE=0.001` means that you get good model but you still don't know what Laplace Loss you get after all. On the other hand `Loss_1=4.321` means lower bound for Laplace Loss (drop some constant for better understanding)."
  },
  "source": "meta"
}