{
  "id": 551533,
  "title": "One trick to optimize thresholds",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/551533",
  "author_name": "Gunes Evitan",
  "post_date": "2024-12-13T18:33:25.279000",
  "votes": 20,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I saw that thresholds are initialized like this <code>[0.5, 1.5, 2.5]</code> on public notebooks. I do it like this on my training pipeline.</p>\n<pre><code>oof_mask = df[].notna()\noof_initial_thresholds = df.loc[oof_mask].groupby()[].mean().iloc[:].values.tolist()\noof_optimized_thresholds = minimize(\n    metrics.rounding_optimization_function,\n    x0=oof_initial_thresholds,\n    args=(df.loc[oof_mask, target], df.loc[oof_mask, ]),\n    method=\n).x\n</code></pre>\n<p>This helps to initialize thresholds dynamically while experimenting and leads to better convergence.</p>",
  "messages": [
    {
      "id": 3071409,
      "postDate": "2024-12-13T18:33:25.280Z",
      "content": "<p>I saw that thresholds are initialized like this <code>[0.5, 1.5, 2.5]</code> on public notebooks. I do it like this on my training pipeline.</p>\n<pre><code>oof_mask = df[].notna()\noof_initial_thresholds = df.loc[oof_mask].groupby()[].mean().iloc[:].values.tolist()\noof_optimized_thresholds = minimize(\n    metrics.rounding_optimization_function,\n    x0=oof_initial_thresholds,\n    args=(df.loc[oof_mask, target], df.loc[oof_mask, ]),\n    method=\n).x\n</code></pre>\n<p>This helps to initialize thresholds dynamically while experimenting and leads to better convergence.</p>",
      "rawMarkdown": "I saw that thresholds are initialized like this `[0.5, 1.5, 2.5]` on public notebooks. I do it like this on my training pipeline.\n\n```python\noof_mask = df['prediction'].notna()\noof_initial_thresholds = df.loc[oof_mask].groupby('target')['prediction'].mean().iloc[1:].values.tolist()\noof_optimized_thresholds = minimize(\n    metrics.rounding_optimization_function,\n    x0=oof_initial_thresholds,\n    args=(df.loc[oof_mask, target], df.loc[oof_mask, 'prediction']),\n    method='Nelder-Mead'\n).x\n```\nThis helps to initialize thresholds dynamically while experimenting and leads to better convergence.",
      "votes": 19
    },
    {
      "id": 3072194,
      "postDate": "2024-12-14T20:46:14.913Z",
      "content": "<p>Thanks for sharing! If I understand this correctly, your initialized thresholds are the average model predictions for each SII category, which I expect to be close to the actual category values. For example the model on average predicts value <code>p1</code> for SII=1, and you take the predictions <code>[p1,p2,p3]</code> as the initial guesses for thresholds. How is this better than, say, <code>(p0+p1)/2, (p1+p2)/2, (p2+p3)/2</code> ?</p>\n<blockquote>\n  <p>leads to better convergence.</p>\n</blockquote>\n<p>How did you measure better convergence? Standard deviation of optimized threshold values across different seeds, or something like that?</p>",
      "rawMarkdown": "Thanks for sharing! If I understand this correctly, your initialized thresholds are the average model predictions for each SII category, which I expect to be close to the actual category values. For example the model on average predicts value `p1` for SII=1, and you take the predictions `[p1,p2,p3]` as the initial guesses for thresholds. How is this better than, say, `(p0+p1)/2, (p1+p2)/2, (p2+p3)/2` ?\n\n>  leads to better convergence.\n\nHow did you measure better convergence? Standard deviation of optimized threshold values across different seeds, or something like that?",
      "votes": 3,
      "replies": [
        {
          "id": 3072421,
          "postDate": "2024-12-15T06:56:35.650Z",
          "content": "<p>Yes, I do exactly what you understood. I haven't tried initializing from <code>(p0+p1)/2, (p1+p2)/2, (p2+p3)/2</code> but I think it's worth trying. Given the search space of 3 thresholds, the local minima is the point where oof qwk is the best and initializing from <code>[0.5, 1.5, 2.5]</code> always stucks at local minima. You can compare by looking at your oof qwk scores.</p>",
          "rawMarkdown": "Yes, I do exactly what you understood. I haven't tried initializing from `(p0+p1)/2, (p1+p2)/2, (p2+p3)/2` but I think it's worth trying. Given the search space of 3 thresholds, the local minima is the point where oof qwk is the best and initializing from `[0.5, 1.5, 2.5]` always stucks at local minima. You can compare by looking at your oof qwk scores.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3072584,
      "postDate": "2024-12-15T12:14:22.497Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3072194,
      "author_name": "Taizhuo Tang",
      "author_url": "",
      "post_date": "2024-12-14T20:46:14.913000",
      "content": "<p>Thanks for sharing! If I understand this correctly, your initialized thresholds are the average model predictions for each SII category, which I expect to be close to the actual category values. For example the model on average predicts value <code>p1</code> for SII=1, and you take the predictions <code>[p1,p2,p3]</code> as the initial guesses for thresholds. How is this better than, say, <code>(p0+p1)/2, (p1+p2)/2, (p2+p3)/2</code> ?</p>\n<blockquote>\n  <p>leads to better convergence.</p>\n</blockquote>\n<p>How did you measure better convergence? Standard deviation of optimized threshold values across different seeds, or something like that?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3072421,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2024-12-15T06:56:35.650000",
          "content": "<p>Yes, I do exactly what you understood. I haven't tried initializing from <code>(p0+p1)/2, (p1+p2)/2, (p2+p3)/2</code> but I think it's worth trying. Given the search space of 3 thresholds, the local minima is the point where oof qwk is the best and initializing from <code>[0.5, 1.5, 2.5]</code> always stucks at local minima. You can compare by looking at your oof qwk scores.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3072584,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-12-15T12:14:22.497000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3071409": "I saw that thresholds are initialized like this `[0.5, 1.5, 2.5]` on public notebooks. I do it like this on my training pipeline.\n\n```python\noof_mask = df['prediction'].notna()\noof_initial_thresholds = df.loc[oof_mask].groupby('target')['prediction'].mean().iloc[1:].values.tolist()\noof_optimized_thresholds = minimize(\n    metrics.rounding_optimization_function,\n    x0=oof_initial_thresholds,\n    args=(df.loc[oof_mask, target], df.loc[oof_mask, 'prediction']),\n    method='Nelder-Mead'\n).x\n```\nThis helps to initialize thresholds dynamically while experimenting and leads to better convergence.",
    "3072194": "Thanks for sharing! If I understand this correctly, your initialized thresholds are the average model predictions for each SII category, which I expect to be close to the actual category values. For example the model on average predicts value `p1` for SII=1, and you take the predictions `[p1,p2,p3]` as the initial guesses for thresholds. How is this better than, say, `(p0+p1)/2, (p1+p2)/2, (p2+p3)/2` ?\n\n>  leads to better convergence.\n\nHow did you measure better convergence? Standard deviation of optimized threshold values across different seeds, or something like that?",
    "3072584": ""
  }
}