{
  "id": 137467,
  "title": "Optimum Clipping",
  "url": "/competitions/deepfake-detection-challenge/discussion/137467",
  "author_name": "",
  "post_date": "2020-03-20T21:36:08.808688500Z",
  "votes": 3,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Any recommendations for determining optimum clipping? </p>\n\n<p>What I mean is guessing 0 when then answer is 1 can lead to almost infinite error because divide by zero. Should I guess 0.05, 0.1, 0.15 instead? Is there a mathematical way to do this or just experiment?</p>",
  "messages": [
    {
      "id": "781056",
      "postDate": "03/20/2020 21:36:08",
      "content": "<p>Any recommendations for determining optimum clipping? </p>\n\n<p>What I mean is guessing 0 when then answer is 1 can lead to almost infinite error because divide by zero. Should I guess 0.05, 0.1, 0.15 instead? Is there a mathematical way to do this or just experiment?</p>",
      "rawMarkdown": "Any recommendations for determining optimum clipping? \n\nWhat I mean is guessing 0 when then answer is 1 can lead to almost infinite error because divide by zero. Should I guess 0.05, 0.1, 0.15 instead? Is there a mathematical way to do this or just experiment?",
      "votes": null
    },
    {
      "id": "781063",
      "postDate": "03/20/2020 21:46:38",
      "content": "<p>It won't lead to infinite error, but still a pretty big number. (It won't divide by zero because they add a small term, 1e-15.)</p>\n\n<p>The problem with clipping is that it also stops really good guesses from lowering the loss. </p>",
      "rawMarkdown": "It won't lead to infinite error, but still a pretty big number. (It won't divide by zero because they add a small term, 1e-15.)\n\nThe problem with clipping is that it also stops really good guesses from lowering the loss.",
      "votes": null
    },
    {
      "id": "781102",
      "postDate": "03/20/2020 23:18:56",
      "content": "<p>Yes that is why I ask for a mathematical way to find optimum clipping</p>",
      "rawMarkdown": "Yes that is why I ask for a mathematical way to find optimum clipping",
      "votes": null
    },
    {
      "id": "781292",
      "postDate": "03/21/2020 05:44:28",
      "content": "<p>You can get where to clip by the accuracy of your model, if the model is getting 90% right on your set and still going all out (making preds like 0.999 or 0.0001), you should clip to (0.85 - 0.95) and (0.15 - 0.05) and higher just clip to (0.01, 0.99)</p>",
      "rawMarkdown": "You can get where to clip by the accuracy of your model, if the model is getting 90% right on your set and still going all out (making preds like 0.999 or 0.0001), you should clip to (0.85 - 0.95) and (0.15 - 0.05) and higher just clip to (0.01, 0.99)",
      "votes": null
    },
    {
      "id": "781433",
      "postDate": "03/21/2020 09:22:44",
      "content": "<p>We're using a simple clipping:\n<code>np.clip(prob, 0.00001, 0.99999)</code></p>",
      "rawMarkdown": "We're using a simple clipping:\n`np.clip(prob, 0.00001, 0.99999)`",
      "votes": null
    },
    {
      "id": "781495",
      "postDate": "03/21/2020 11:12:30",
      "content": "<p>I wonder what would be the accuracy of your model, that you letting it go almost all out, 0.00001 is enough to destroy when it is not accurate.</p>",
      "rawMarkdown": "I wonder what would be the accuracy of your model, that you letting it go almost all out, 0.00001 is enough to destroy when it is not accurate.",
      "votes": null
    },
    {
      "id": "782754",
      "postDate": "03/22/2020 16:08:23",
      "content": "<p>I would suggest looking at how clipping affects your loss on your own validation/test set and use the tightest clipping that doesn't increase your loss.  For example, if you find that you get roughly the same loss for clipping at .99/.01 as .98/.02, then I would clip at .98/.02.   And if you want to use up submissions with these types of tests, you could experiment from there with tightening further.</p>",
      "rawMarkdown": "I would suggest looking at how clipping affects your loss on your own validation/test set and use the tightest clipping that doesn't increase your loss.  For example, if you find that you get roughly the same loss for clipping at .99/.01 as .98/.02, then I would clip at .98/.02.   And if you want to use up submissions with these types of tests, you could experiment from there with tightening further.",
      "votes": null
    },
    {
      "id": "784546",
      "postDate": "03/24/2020 09:59:28",
      "content": "<p>Well, I wouldn't rely on mathematics, but rather on experimentation. The LB score and the validation scores typically vary quite a bit. In my case, a good validation clipping was around 0.01-0.99 and 0.1-0.9 was slightly worse. For the submitted LB score, it was the exact opposite, probably due to more misinterpreted videos.</p>\n\n<p>In any case I recommend validating first to 0.1/0.9 to know if the model is OK and then to something like 0.01/0.99 when the model proves to be fairly accurate. Beyond that, the gains are really marginal and the cost of errors skyrocket.</p>",
      "rawMarkdown": "Well, I wouldn't rely on mathematics, but rather on experimentation. The LB score and the validation scores typically vary quite a bit. In my case, a good validation clipping was around 0.01-0.99 and 0.1-0.9 was slightly worse. For the submitted LB score, it was the exact opposite, probably due to more misinterpreted videos.\n\nIn any case I recommend validating first to 0.1/0.9 to know if the model is OK and then to something like 0.01/0.99 when the model proves to be fairly accurate. Beyond that, the gains are really marginal and the cost of errors skyrocket.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 781063,
      "author_name": "humananalog",
      "author_url": "",
      "post_date": "03/20/2020 21:46:38",
      "content": "<p>It won't lead to infinite error, but still a pretty big number. (It won't divide by zero because they add a small term, 1e-15.)</p>\n\n<p>The problem with clipping is that it also stops really good guesses from lowering the loss. </p>",
      "votes": null,
      "replies": [
        {
          "id": 781102,
          "author_name": "sethkitchen",
          "author_url": "",
          "post_date": "03/20/2020 23:18:56",
          "content": "<p>Yes that is why I ask for a mathematical way to find optimum clipping</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 781292,
      "author_name": "harshitsheoran",
      "author_url": "",
      "post_date": "03/21/2020 05:44:28",
      "content": "<p>You can get where to clip by the accuracy of your model, if the model is getting 90% right on your set and still going all out (making preds like 0.999 or 0.0001), you should clip to (0.85 - 0.95) and (0.15 - 0.05) and higher just clip to (0.01, 0.99)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 781433,
      "author_name": "mpware",
      "author_url": "",
      "post_date": "03/21/2020 09:22:44",
      "content": "<p>We're using a simple clipping:\n<code>np.clip(prob, 0.00001, 0.99999)</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 781495,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "03/21/2020 11:12:30",
          "content": "<p>I wonder what would be the accuracy of your model, that you letting it go almost all out, 0.00001 is enough to destroy when it is not accurate.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 782754,
      "author_name": "maierman",
      "author_url": "",
      "post_date": "03/22/2020 16:08:23",
      "content": "<p>I would suggest looking at how clipping affects your loss on your own validation/test set and use the tightest clipping that doesn't increase your loss.  For example, if you find that you get roughly the same loss for clipping at .99/.01 as .98/.02, then I would clip at .98/.02.   And if you want to use up submissions with these types of tests, you could experiment from there with tightening further.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 784546,
      "author_name": "dagnelies",
      "author_url": "",
      "post_date": "03/24/2020 09:59:28",
      "content": "<p>Well, I wouldn't rely on mathematics, but rather on experimentation. The LB score and the validation scores typically vary quite a bit. In my case, a good validation clipping was around 0.01-0.99 and 0.1-0.9 was slightly worse. For the submitted LB score, it was the exact opposite, probably due to more misinterpreted videos.</p>\n\n<p>In any case I recommend validating first to 0.1/0.9 to know if the model is OK and then to something like 0.01/0.99 when the model proves to be fairly accurate. Beyond that, the gains are really marginal and the cost of errors skyrocket.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "781056": "Any recommendations for determining optimum clipping? \n\nWhat I mean is guessing 0 when then answer is 1 can lead to almost infinite error because divide by zero. Should I guess 0.05, 0.1, 0.15 instead? Is there a mathematical way to do this or just experiment?",
    "781063": "It won't lead to infinite error, but still a pretty big number. (It won't divide by zero because they add a small term, 1e-15.)\n\nThe problem with clipping is that it also stops really good guesses from lowering the loss.",
    "781102": "Yes that is why I ask for a mathematical way to find optimum clipping",
    "781292": "You can get where to clip by the accuracy of your model, if the model is getting 90% right on your set and still going all out (making preds like 0.999 or 0.0001), you should clip to (0.85 - 0.95) and (0.15 - 0.05) and higher just clip to (0.01, 0.99)",
    "781433": "We're using a simple clipping:\n`np.clip(prob, 0.00001, 0.99999)`",
    "781495": "I wonder what would be the accuracy of your model, that you letting it go almost all out, 0.00001 is enough to destroy when it is not accurate.",
    "782754": "I would suggest looking at how clipping affects your loss on your own validation/test set and use the tightest clipping that doesn't increase your loss.  For example, if you find that you get roughly the same loss for clipping at .99/.01 as .98/.02, then I would clip at .98/.02.   And if you want to use up submissions with these types of tests, you could experiment from there with tightening further.",
    "784546": "Well, I wouldn't rely on mathematics, but rather on experimentation. The LB score and the validation scores typically vary quite a bit. In my case, a good validation clipping was around 0.01-0.99 and 0.1-0.9 was slightly worse. For the submitted LB score, it was the exact opposite, probably due to more misinterpreted videos.\n\nIn any case I recommend validating first to 0.1/0.9 to know if the model is OK and then to something like 0.01/0.99 when the model proves to be fairly accurate. Beyond that, the gains are really marginal and the cost of errors skyrocket."
  },
  "source": "meta"
}