{
  "id": 552090,
  "title": "Threshold calculation in the best public notebooks",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/552090",
  "author_name": "Vladimir Demidov",
  "post_date": "2024-12-17T15:19:10.981000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I'm wondering If anybody tried to incorporate the way of the thresholds calculation from <a href=\"https://www.kaggle.com/code/vitalykudelya/cmi-issues-with-the-metric-and-baseline\" target=\"_blank\">this</a> 0.495 notebook into your solution, especially if you are working on  improving the best public 0.497 notebook? As you see, <a href=\"https://www.kaggle.com/vitalykudelya\" target=\"_blank\">@vitalykudelya</a> presents great insight about the vulnerability of the standard way but does not change the further pipeline for some reason; the author of the best 0.497 public notebook is missing this detail as well.</p>\n<p>So I experiment by myself in my <a href=\"https://www.kaggle.com/code/yekenot/cmi-piu-deeptables-nn-cv-0-482\" target=\"_blank\">notebook</a>. I found that new thresholds providing global maximums lead to an improvement in CV of around 0.01 and LB of 0.005 if the first standard way was semi-successful. In my current best version of the notebook, the first step of optimization didn't almost anything (it fell down to unlucky local maximums). But the second search of thresholds has calculated values for the global maximums, and CV rose up drastically. It's just an interesting observation; maybe it will be helpful for you.</p>",
  "messages": [
    {
      "id": 3074336,
      "postDate": "2024-12-17T15:19:10.980Z",
      "content": "<p>I'm wondering If anybody tried to incorporate the way of the thresholds calculation from <a href=\"https://www.kaggle.com/code/vitalykudelya/cmi-issues-with-the-metric-and-baseline\" target=\"_blank\">this</a> 0.495 notebook into your solution, especially if you are working on  improving the best public 0.497 notebook? As you see, <a href=\"https://www.kaggle.com/vitalykudelya\" target=\"_blank\">@vitalykudelya</a> presents great insight about the vulnerability of the standard way but does not change the further pipeline for some reason; the author of the best 0.497 public notebook is missing this detail as well.</p>\n<p>So I experiment by myself in my <a href=\"https://www.kaggle.com/code/yekenot/cmi-piu-deeptables-nn-cv-0-482\" target=\"_blank\">notebook</a>. I found that new thresholds providing global maximums lead to an improvement in CV of around 0.01 and LB of 0.005 if the first standard way was semi-successful. In my current best version of the notebook, the first step of optimization didn't almost anything (it fell down to unlucky local maximums). But the second search of thresholds has calculated values for the global maximums, and CV rose up drastically. It's just an interesting observation; maybe it will be helpful for you.</p>",
      "rawMarkdown": "I'm wondering If anybody tried to incorporate the way of the thresholds calculation from [this](https://www.kaggle.com/code/vitalykudelya/cmi-issues-with-the-metric-and-baseline) 0.495 notebook into your solution, especially if you are working on ~~overfitting~~ improving the best public 0.497 notebook? As you see, @vitalykudelya presents great insight about the vulnerability of the standard way but does not change the further pipeline for some reason; the author of the best 0.497 public notebook is missing this detail as well.\n\nSo I experiment by myself in my [notebook](https://www.kaggle.com/code/yekenot/cmi-piu-deeptables-nn-cv-0-482). I found that new thresholds providing global maximums lead to an improvement in CV of around 0.01 and LB of 0.005 if the first standard way was semi-successful. In my current best version of the notebook, the first step of optimization didn't almost anything (it fell down to unlucky local maximums). But the second search of thresholds has calculated values for the global maximums, and CV rose up drastically. It's just an interesting observation; maybe it will be helpful for you.",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3074336": "I'm wondering If anybody tried to incorporate the way of the thresholds calculation from [this](https://www.kaggle.com/code/vitalykudelya/cmi-issues-with-the-metric-and-baseline) 0.495 notebook into your solution, especially if you are working on ~~overfitting~~ improving the best public 0.497 notebook? As you see, @vitalykudelya presents great insight about the vulnerability of the standard way but does not change the further pipeline for some reason; the author of the best 0.497 public notebook is missing this detail as well.\n\nSo I experiment by myself in my [notebook](https://www.kaggle.com/code/yekenot/cmi-piu-deeptables-nn-cv-0-482). I found that new thresholds providing global maximums lead to an improvement in CV of around 0.01 and LB of 0.005 if the first standard way was semi-successful. In my current best version of the notebook, the first step of optimization didn't almost anything (it fell down to unlucky local maximums). But the second search of thresholds has calculated values for the global maximums, and CV rose up drastically. It's just an interesting observation; maybe it will be helpful for you."
  }
}