{
  "id": 302130,
  "title": "Here is the magic: Estimating recall and precision from your public LB score",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/302130",
  "author_name": "Bilzard",
  "post_date": "2022-01-21T01:22:21.338000",
  "votes": 23,
  "comment_count": 11,
  "views": 0,
  "content": "<h1>To be short</h1>\n<p>We can estimate recall and precision of your prediction model on public LB data.<br>\nWhat we need is:</p>\n<ol>\n<li>Randomly dropping your model's prediction probability 1 - rho</li>\n<li>You need at least two submissions of different dropping probability (say, rho=1 and 0.5) and observe F2</li>\n<li>You can calculate precision and recall using the below code</li>\n</ol>\n<h1>Code</h1>\n<pre><code>def drop_pred(bboxes, confs, p_keep):\n    '''\n    randomly drop prediction for the probability (1 - p_keep).\n    '''\n    if p_keep == 1:\n        return bboxes, confs\n    bboxes = bboxes.copy()\n    confs = confs.copy()\n    assert len(bboxes) == len(confs)\n    pp = np.random.uniform(size=len(bboxes))\n    bboxes = bboxes[pp &lt;= p_keep]\n    confs = confs[pp &lt;= p_keep]\n    return bboxes, confs\n</code></pre>\n<pre><code>def estimate_recall_and_precision(f2, f2s, rhos):\n    recalls, precisions = [], []\n    for f2_rho, rho in zip(f2s, rhos):\n        fn_tp = 5 / 4 * rho / (1 - rho) * (1 / f2_rho - 1 / f2) - 1\n        fp_tp = 5 / (1 - rho) * (1 / f2 - 1 / f2_rho * rho) - 1\n        recall = 1 / (fn_tp + 1)\n        precision = 1 / (fp_tp + 1)\n        recalls.append(recall)\n        precisions.append(precision)\n\n    return recalls, precisions\n</code></pre>\n<h1>Detail</h1>\n<p>It was somehow tedious to verify, it might be some tiny mistake.<br>\nBut I tested the final equation by simulation, so I think the code above is credible.</p>\n<hr>\n<p>The F2 score is calculated by equation below:</p>\n<p>$$<br>\nF_2 = \\frac{5TP}{5TP+4FN+FP}<br>\n$$</p>\n<p>If we drop 1 - rho predictions, TP, FP, and FN changes to </p>\n<p>$$<br>\n\\begin{eqnarray}<br>\nTP^\\prime &amp;=&amp; \\rho TP \\\\<br>\nFN^\\prime &amp;=&amp; FN + (1 - \\rho) TP \\\\<br>\nFP^\\prime &amp;=&amp; \\rho FP<br>\n\\end{eqnarray}<br>\n$$</p>\n<p>And the new F2 score F2_rho is calculated as:</p>\n<p>$$<br>\nF_2^\\rho = \\frac{5\\rho TP}{5\\rho TP+4(FN+(1-\\rho)TP)+\\rho FP}<br>\n$$</p>\n<p>Clearing out denominator and we get the equation of FN/TP and FP/TP:</p>\n<p>$$<br>\n4 \\frac{FN}{TP} + \\rho \\frac{FP}{TP} = \\rho \\left( \\frac{5}{F_2^\\rho} - 1 \\right) - 4<br>\n$$</p>\n<p>Because the equation has two variables, we need at least two observation of F2_rho based on the different dropout ratio 1 - rho. Given we already have observation F2 (without dropout), we only require another observation of F_2_rho (e.g, rho=0.5).</p>\n<p>Then we have two linear equations:</p>\n<p>$$<br>\n4 \\frac{FN}{TP} + \\frac{FP}{TP} = \\left( \\frac{5}{F_2} - 1 \\right) - 4 \\tag{a}<br>\n$$<br>\n$$<br>\n4 \\frac{FN}{TP} + \\rho \\frac{FP}{TP} = \\rho \\left( \\frac{5}{F_2^\\rho} - 1 \\right) - 4 \\tag{b}<br>\n$$</p>\n<p>Solving above equations, we get</p>\n<p>$$<br>\n\\frac{FN}{TP}=\\frac{5}{4}\\frac{\\rho}{1-\\rho}\\left(\\frac{1}{F_2^{\\rho}}-\\frac{1}{F_2}\\right) - 1 \\tag{1}<br>\n$$</p>\n<p>$$<br>\n\\frac{FP}{TP}=\\frac{5}{1-\\rho}\\left( \\frac{1}{F_2} - \\frac{\\rho}{F_2^{\\rho}} \\right) -1 \\tag{2}<br>\n$$</p>\n<p>And since recall=TP/(TP+FN) and precision=TP/(TP+FP), we can calculate recall and precision from this ratio.</p>\n<p><a href=\"https://ibb.co/XpgvcGT\"><img src=\"https://i.ibb.co/3WqLgPH/Screen-Shot-2022-01-21-at-9-36-00.png\" alt=\"Screen-Shot-2022-01-21-at-9-36-00\"></a></p>\n<h1>Reference</h1>\n<p>This idea was inspired from below discussion by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>:<br>\n<a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300405#1656348\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300405#1656348</a></p>\n<hr>\n<h1>Update Note</h1>\n<ul>\n<li>2022/1/21 19:32 JST refactor equation &amp; code: use more simpler form. Also do a simple degrade check by feeding few experiment data</li>\n<li>2022/1/22 18:19 JST simplify equation</li>\n</ul>",
  "messages": [
    {
      "id": 1658380,
      "postDate": "2022-01-21T01:22:21.340Z",
      "content": "<h1>To be short</h1>\n<p>We can estimate recall and precision of your prediction model on public LB data.<br>\nWhat we need is:</p>\n<ol>\n<li>Randomly dropping your model's prediction probability 1 - rho</li>\n<li>You need at least two submissions of different dropping probability (say, rho=1 and 0.5) and observe F2</li>\n<li>You can calculate precision and recall using the below code</li>\n</ol>\n<h1>Code</h1>\n<pre><code>def drop_pred(bboxes, confs, p_keep):\n    '''\n    randomly drop prediction for the probability (1 - p_keep).\n    '''\n    if p_keep == 1:\n        return bboxes, confs\n    bboxes = bboxes.copy()\n    confs = confs.copy()\n    assert len(bboxes) == len(confs)\n    pp = np.random.uniform(size=len(bboxes))\n    bboxes = bboxes[pp &lt;= p_keep]\n    confs = confs[pp &lt;= p_keep]\n    return bboxes, confs\n</code></pre>\n<pre><code>def estimate_recall_and_precision(f2, f2s, rhos):\n    recalls, precisions = [], []\n    for f2_rho, rho in zip(f2s, rhos):\n        fn_tp = 5 / 4 * rho / (1 - rho) * (1 / f2_rho - 1 / f2) - 1\n        fp_tp = 5 / (1 - rho) * (1 / f2 - 1 / f2_rho * rho) - 1\n        recall = 1 / (fn_tp + 1)\n        precision = 1 / (fp_tp + 1)\n        recalls.append(recall)\n        precisions.append(precision)\n\n    return recalls, precisions\n</code></pre>\n<h1>Detail</h1>\n<p>It was somehow tedious to verify, it might be some tiny mistake.<br>\nBut I tested the final equation by simulation, so I think the code above is credible.</p>\n<hr>\n<p>The F2 score is calculated by equation below:</p>\n<p>$$<br>\nF_2 = \\frac{5TP}{5TP+4FN+FP}<br>\n$$</p>\n<p>If we drop 1 - rho predictions, TP, FP, and FN changes to </p>\n<p>$$<br>\n\\begin{eqnarray}<br>\nTP^\\prime &amp;=&amp; \\rho TP \\\\<br>\nFN^\\prime &amp;=&amp; FN + (1 - \\rho) TP \\\\<br>\nFP^\\prime &amp;=&amp; \\rho FP<br>\n\\end{eqnarray}<br>\n$$</p>\n<p>And the new F2 score F2_rho is calculated as:</p>\n<p>$$<br>\nF_2^\\rho = \\frac{5\\rho TP}{5\\rho TP+4(FN+(1-\\rho)TP)+\\rho FP}<br>\n$$</p>\n<p>Clearing out denominator and we get the equation of FN/TP and FP/TP:</p>\n<p>$$<br>\n4 \\frac{FN}{TP} + \\rho \\frac{FP}{TP} = \\rho \\left( \\frac{5}{F_2^\\rho} - 1 \\right) - 4<br>\n$$</p>\n<p>Because the equation has two variables, we need at least two observation of F2_rho based on the different dropout ratio 1 - rho. Given we already have observation F2 (without dropout), we only require another observation of F_2_rho (e.g, rho=0.5).</p>\n<p>Then we have two linear equations:</p>\n<p>$$<br>\n4 \\frac{FN}{TP} + \\frac{FP}{TP} = \\left( \\frac{5}{F_2} - 1 \\right) - 4 \\tag{a}<br>\n$$<br>\n$$<br>\n4 \\frac{FN}{TP} + \\rho \\frac{FP}{TP} = \\rho \\left( \\frac{5}{F_2^\\rho} - 1 \\right) - 4 \\tag{b}<br>\n$$</p>\n<p>Solving above equations, we get</p>\n<p>$$<br>\n\\frac{FN}{TP}=\\frac{5}{4}\\frac{\\rho}{1-\\rho}\\left(\\frac{1}{F_2^{\\rho}}-\\frac{1}{F_2}\\right) - 1 \\tag{1}<br>\n$$</p>\n<p>$$<br>\n\\frac{FP}{TP}=\\frac{5}{1-\\rho}\\left( \\frac{1}{F_2} - \\frac{\\rho}{F_2^{\\rho}} \\right) -1 \\tag{2}<br>\n$$</p>\n<p>And since recall=TP/(TP+FN) and precision=TP/(TP+FP), we can calculate recall and precision from this ratio.</p>\n<p><a href=\"https://ibb.co/XpgvcGT\"><img src=\"https://i.ibb.co/3WqLgPH/Screen-Shot-2022-01-21-at-9-36-00.png\" alt=\"Screen-Shot-2022-01-21-at-9-36-00\"></a></p>\n<h1>Reference</h1>\n<p>This idea was inspired from below discussion by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>:<br>\n<a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300405#1656348\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300405#1656348</a></p>\n<hr>\n<h1>Update Note</h1>\n<ul>\n<li>2022/1/21 19:32 JST refactor equation &amp; code: use more simpler form. Also do a simple degrade check by feeding few experiment data</li>\n<li>2022/1/22 18:19 JST simplify equation</li>\n</ul>",
      "rawMarkdown": "# To be short\n\nWe can estimate recall and precision of your prediction model on public LB data.\nWhat we need is:\n\n1. Randomly dropping your model's prediction probability 1 - rho\n2. You need at least two submissions of different dropping probability (say, rho=1 and 0.5) and observe F2\n3. You can calculate precision and recall using the below code\n\n# Code \n\n```\ndef drop_pred(bboxes, confs, p_keep):\n    '''\n    randomly drop prediction for the probability (1 - p_keep).\n    '''\n    if p_keep == 1:\n        return bboxes, confs\n    bboxes = bboxes.copy()\n    confs = confs.copy()\n    assert len(bboxes) == len(confs)\n    pp = np.random.uniform(size=len(bboxes))\n    bboxes = bboxes[pp <= p_keep]\n    confs = confs[pp <= p_keep]\n    return bboxes, confs\n```\n\n```python\ndef estimate_recall_and_precision(f2, f2s, rhos):\n    recalls, precisions = [], []\n    for f2_rho, rho in zip(f2s, rhos):\n        fn_tp = 5 / 4 * rho / (1 - rho) * (1 / f2_rho - 1 / f2) - 1\n        fp_tp = 5 / (1 - rho) * (1 / f2 - 1 / f2_rho * rho) - 1\n        recall = 1 / (fn_tp + 1)\n        precision = 1 / (fp_tp + 1)\n        recalls.append(recall)\n        precisions.append(precision)\n\n    return recalls, precisions\n```\n\n# Detail\n\nIt was somehow tedious to verify, it might be some tiny mistake.\nBut I tested the final equation by simulation, so I think the code above is credible.\n\n---\n\nThe F2 score is calculated by equation below:\n\n$$\nF_2 = \\frac{5TP}{5TP+4FN+FP}\n$$\n\nIf we drop 1 - rho predictions, TP, FP, and FN changes to \n\n$$\n\\begin{eqnarray}\nTP^\\prime &=& \\rho TP \\\\\\\\\nFN^\\prime &=& FN + (1 - \\rho) TP \\\\\\\\\nFP^\\prime &=& \\rho FP\n\\end{eqnarray}\n$$\n\nAnd the new F2 score F2_rho is calculated as:\n\n$$\nF_2^\\rho = \\frac{5\\rho TP}{5\\rho TP+4(FN+(1-\\rho)TP)+\\rho FP}\n$$\n\nClearing out denominator and we get the equation of FN/TP and FP/TP:\n\n$$\n4 \\frac{FN}{TP} + \\rho \\frac{FP}{TP} = \\rho \\left( \\frac{5}{F_2^\\rho} - 1 \\right) - 4\n$$\n\nBecause the equation has two variables, we need at least two observation of F2_rho based on the different dropout ratio 1 - rho. Given we already have observation F2 (without dropout), we only require another observation of F_2_rho (e.g, rho=0.5).\n\nThen we have two linear equations:\n\n$$\n4 \\frac{FN}{TP} + \\frac{FP}{TP} = \\left( \\frac{5}{F_2} - 1 \\right) - 4 \\tag{a}\n$$\n$$\n4 \\frac{FN}{TP} + \\rho \\frac{FP}{TP} = \\rho \\left( \\frac{5}{F_2^\\rho} - 1 \\right) - 4 \\tag{b}\n$$\n\nSolving above equations, we get\n\n$$\n\\frac{FN}{TP}=\\frac{5}{4}\\frac{\\rho}{1-\\rho}\\left(\\frac{1}{F_2^{\\rho}}-\\frac{1}{F_2}\\right) - 1 \\tag{1}\n$$\n\n$$\n\\frac{FP}{TP}=\\frac{5}{1-\\rho}\\left( \\frac{1}{F_2} - \\frac{\\rho}{F_2^{\\rho}} \\right) -1 \\tag{2}\n$$\n\nAnd since recall=TP/(TP+FN) and precision=TP/(TP+FP), we can calculate recall and precision from this ratio.\n\n<a href=\"https://ibb.co/XpgvcGT\"><img src=\"https://i.ibb.co/3WqLgPH/Screen-Shot-2022-01-21-at-9-36-00.png\" alt=\"Screen-Shot-2022-01-21-at-9-36-00\" border=\"0\"></a>\n\n# Reference\n\nThis idea was inspired from below discussion by @hengck23:\nhttps://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300405#1656348\n\n----\n\n# Update Note\n\n* 2022/1/21 19:32 JST refactor equation & code: use more simpler form. Also do a simple degrade check by feeding few experiment data\n* 2022/1/22 18:19 JST simplify equation",
      "votes": 23
    },
    {
      "id": 1658835,
      "postDate": "2022-01-21T10:49:29.110Z",
      "content": "<p>make sure you have lots of background images (and many variations, e.g fish, bubble …) to compute your model precision in validation.</p>\n<p>by comparing public test precision (or better still, comparing fp rate), we can guess the quality and number of background images.</p>\n<p>top kagglers are good are modeling deep network and modeling unseen data. modeling unseen data is competition exploitation and is the key to winning. (expert kagglers don't experience shakeup)</p>\n<p>actually just private and public set is not enough for competition design to prevent exploitation.</p>",
      "rawMarkdown": "make sure you have lots of background images (and many variations, e.g fish, bubble ...) to compute your model precision in validation.\n\nby comparing public test precision (or better still, comparing fp rate), we can guess the quality and number of background images.\n\ntop kagglers are good are modeling deep network and modeling unseen data. modeling unseen data is competition exploitation and is the key to winning. (expert kagglers don't experience shakeup)\n\nactually just private and public set is not enough for competition design to prevent exploitation.",
      "votes": 3,
      "replies": [
        {
          "id": 1658841,
          "postDate": "2022-01-21T10:56:25.233Z",
          "content": "<blockquote>\n  <p>modeling unseen data is competition exploitation and is the key to winning</p>\n</blockquote>\n<p>Thanks for the words. I'll do my best.</p>",
          "rawMarkdown": "> modeling unseen data is competition exploitation and is the key to winning\n\nThanks for the words. I'll do my best."
        }
      ]
    },
    {
      "id": 1658433,
      "postDate": "2022-01-21T03:02:47.020Z",
      "content": "<p>this is human-based meta learning</p>",
      "rawMarkdown": "this is human-based meta learning",
      "votes": 3
    },
    {
      "id": 1658385,
      "postDate": "2022-01-21T01:30:21.610Z",
      "content": "<p>I also make 4 submission of rho=(1.0, 0.7, 0.5, 0.3), and calculated the result.<br>\nI observed high variance std=~0.1 on precision, but the variance of recall is relatively low (std=~0.03).</p>\n<pre><code>         R      P\nrho              \n0.7  0.590  0.424\n0.5  0.560  0.501\n0.3  0.522  0.678\n========================================\n      R     P\nmean: 0.557 0.534\nstd:  0.028 0.106\n</code></pre>",
      "rawMarkdown": "I also make 4 submission of rho=(1.0, 0.7, 0.5, 0.3), and calculated the result.\nI observed high variance std=~0.1 on precision, but the variance of recall is relatively low (std=~0.03).\n\n```\n         R      P\nrho              \n0.7  0.590  0.424\n0.5  0.560  0.501\n0.3  0.522  0.678\n========================================\n      R     P\nmean: 0.557 0.534\nstd:  0.028 0.106\n```",
      "votes": 1,
      "replies": [
        {
          "id": 1658424,
          "postDate": "2022-01-21T02:47:10.193Z",
          "content": "<p>computation result of another 4 submission of different model:<br>\nWe have relatively small variance both on Recall and Precision.</p>\n<pre><code>         R      P\nrho\n0.7  0.614  0.554\n0.5  0.631  0.506\n0.3  0.625  0.521\n========================================\n      R     P\nmean: 0.623 0.527\nstd:  0.007 0.020\n</code></pre>",
          "rawMarkdown": "computation result of another 4 submission of different model:\nWe have relatively small variance both on Recall and Precision.\n\n```\n         R      P\nrho\n0.7  0.614  0.554\n0.5  0.631  0.506\n0.3  0.625  0.521\n========================================\n      R     P\nmean: 0.623 0.527\nstd:  0.007 0.020\n```",
          "votes": 1
        },
        {
          "id": 1658520,
          "postDate": "2022-01-21T05:14:27.773Z",
          "content": "<p>Intuitive comprehension of why variance of precision is higher than that of recall is here.<br>\nSince F2 metrics focus more on recall than precision, our computation is more sensitive to recall than precision. Therefore we have higher variance in precision than recall.<br>\nOf course we need more math to prove this.</p>",
          "rawMarkdown": "Intuitive comprehension of why variance of precision is higher than that of recall is here.\nSince F2 metrics focus more on recall than precision, our computation is more sensitive to recall than precision. Therefore we have higher variance in precision than recall.\nOf course we need more math to prove this."
        },
        {
          "id": 1661221,
          "postDate": "2022-01-23T10:11:31.180Z",
          "content": "<p>Solving the problem with least square fit, I have the following result:</p>\n<p>first model:</p>\n<pre><code>recall = 0.533\nprecision = 0.657\n</code></pre>\n<p>second model:</p>\n<pre><code>recall = 0.626\nprecision = 0.515\n</code></pre>\n<p>This time, we get more higher estimation of precision and slightly less recall in the first model.<br>\nFor the second model, estimated values are almost the same as for the former experiment.</p>",
          "rawMarkdown": "Solving the problem with least square fit, I have the following result:\n\nfirst model:\n```\nrecall = 0.533\nprecision = 0.657\n```\n\nsecond model:\n```\nrecall = 0.626\nprecision = 0.515\n```\n\nThis time, we get more higher estimation of precision and slightly less recall in the first model.\nFor the second model, estimated values are almost the same as for the former experiment."
        }
      ]
    },
    {
      "id": 1658384,
      "postDate": "2022-01-21T01:28:55.097Z",
      "content": "<h1>Simulation result</h1>\n<p>Assumptions:</p>\n<ul>\n<li>F2 score is rounded at decimal point 3</li>\n</ul>\n<p><a href=\"https://ibb.co/qCNVR9J\"><img src=\"https://i.ibb.co/xY1dMGF/Screen-Shot-2022-01-27-at-12-48-39.png\" alt=\"Screen-Shot-2022-01-27-at-12-48-39\"></a></p>",
      "rawMarkdown": "# Simulation result\n\nAssumptions:\n* F2 score is rounded at decimal point 3\n\n<a href=\"https://ibb.co/qCNVR9J\"><img src=\"https://i.ibb.co/xY1dMGF/Screen-Shot-2022-01-27-at-12-48-39.png\" alt=\"Screen-Shot-2022-01-27-at-12-48-39\" border=\"0\"></a>",
      "votes": 1
    },
    {
      "id": 1661207,
      "postDate": "2022-01-23T10:04:53.730Z",
      "content": "<h2>Tips: Estimate Recall and Precision with Least Square Method</h2>\n<p>If we have more than two observation of F2, we can solve this problem as least square regression.<br>\nGiven we obtain N equations (a) for different rho (rho_1, rho_2, …, rho_N), we have equations Ax = b where</p>\n<p>$$<br>\nA = \\begin{bmatrix}<br>\n4 &amp; \\rho_1 \\\\<br>\n4 &amp; \\rho_2 \\\\<br>\n… \\\\<br>\n4 &amp; \\rho_N \\\\<br>\n\\end{bmatrix},<br>\nb = \\begin{bmatrix}<br>\n\\rho_1 (5 / F_2^{\\rho_1} - 1) -4 \\\\<br>\n\\rho_2 (5 / F_2^{\\rho_2} - 1) -4 \\\\<br>\n… \\\\<br>\n\\rho_N (5 / F_2^{\\rho_N} - 1) -4 \\\\<br>\n\\end{bmatrix}, <br>\nx = \\begin{bmatrix}<br>\n\\frac{FN}{TP} \\\\<br>\n\\frac{FP}{TP} \\\\<br>\n\\end{bmatrix}<br>\n$$</p>\n<p>This is an over-determined system, so we can solve this problem by least-square regression[1].</p>\n<p>$$<br>\nmin_x\\|Ax - b\\|<br>\n$$</p>\n<p>The solution is</p>\n<p>$$<br>\nx = (A^\\top A)^{-1}A^\\top b<br>\n$$</p>\n<p>[1] <a href=\"https://en.wikipedia.org/wiki/Overdetermined_system#Approximate_solutions\" target=\"_blank\">https://en.wikipedia.org/wiki/Overdetermined_system#Approximate_solutions</a></p>",
      "rawMarkdown": "## Tips: Estimate Recall and Precision with Least Square Method\n\nIf we have more than two observation of F2, we can solve this problem as least square regression.\nGiven we obtain N equations (a) for different rho (rho_1, rho_2, ..., rho_N), we have equations Ax = b where\n\n$$\nA = \\begin{bmatrix}\n4 & \\rho_1 \\\\\\\\\n4 & \\rho_2 \\\\\\\\\n... \\\\\\\\\n4 & \\rho_N \\\\\\\\\n\\end{bmatrix},\nb = \\begin{bmatrix}\n\\rho_1 (5 / F_2^{\\rho_1} - 1) -4 \\\\\\\\\n\\rho_2 (5 / F_2^{\\rho_2} - 1) -4 \\\\\\\\\n... \\\\\\\\\n\\rho_N (5 / F_2^{\\rho_N} - 1) -4 \\\\\\\\\n\\end{bmatrix}, \nx = \\begin{bmatrix}\n\\frac{FN}{TP} \\\\\\\\\n\\frac{FP}{TP} \\\\\\\\\n\\end{bmatrix}\n$$\n\nThis is an over-determined system, so we can solve this problem by least-square regression[1].\n\n$$\nmin_x\\\\|Ax - b\\\\|\n$$\n\nThe solution is\n\n$$\nx = (A^\\top A)^{-1}A^\\top b\n$$\n\n[1] https://en.wikipedia.org/wiki/Overdetermined_system#Approximate_solutions"
    },
    {
      "id": 1658440,
      "postDate": "2022-01-21T03:09:52.027Z",
      "content": "<h2>Eg. estimating recall &amp; precision for the different train/test scales</h2>\n<p>I estimated recall and precision in the Leader board, and have result below.<br>\nIn this result, recall is higher when we infer larger scale than that in train, whereas precision is nearly the same.<br>\nOf course we need more sample because it depends on the model.<br>\nI don't confirm this result is generally applicable.</p>\n<p>train scale: x2.00, infer scale: x2.00</p>\n<pre><code>      R     P\nmean: 0.557 0.534\nstd:  0.028 0.106\n</code></pre>\n<p>train scale: x2.00, infer scale: x3.20 (x1.60 larger than train scale)</p>\n<pre><code>      R     P\nmean: 0.623 0.527\nstd:  0.007 0.020\n</code></pre>",
      "rawMarkdown": "## Eg. estimating recall & precision for the different train/test scales\n\nI estimated recall and precision in the Leader board, and have result below.\nIn this result, recall is higher when we infer larger scale than that in train, whereas precision is nearly the same.\nOf course we need more sample because it depends on the model.\nI don't confirm this result is generally applicable.\n\ntrain scale: x2.00, infer scale: x2.00\n```\n      R     P\nmean: 0.557 0.534\nstd:  0.028 0.106\n```\n\ntrain scale: x2.00, infer scale: x3.20 (x1.60 larger than train scale)\n```\n      R     P\nmean: 0.623 0.527\nstd:  0.007 0.020\n```"
    },
    {
      "id": 1658827,
      "postDate": "2022-01-21T10:44:31.367Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1658835,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-01-21T10:49:29.110000",
      "content": "<p>make sure you have lots of background images (and many variations, e.g fish, bubble …) to compute your model precision in validation.</p>\n<p>by comparing public test precision (or better still, comparing fp rate), we can guess the quality and number of background images.</p>\n<p>top kagglers are good are modeling deep network and modeling unseen data. modeling unseen data is competition exploitation and is the key to winning. (expert kagglers don't experience shakeup)</p>\n<p>actually just private and public set is not enough for competition design to prevent exploitation.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1658841,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-01-21T10:56:25.233000",
          "content": "<blockquote>\n  <p>modeling unseen data is competition exploitation and is the key to winning</p>\n</blockquote>\n<p>Thanks for the words. I'll do my best.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1658433,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-01-21T03:02:47.020000",
      "content": "<p>this is human-based meta learning</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1658385,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2022-01-21T01:30:21.610000",
      "content": "<p>I also make 4 submission of rho=(1.0, 0.7, 0.5, 0.3), and calculated the result.<br>\nI observed high variance std=~0.1 on precision, but the variance of recall is relatively low (std=~0.03).</p>\n<pre><code>         R      P\nrho              \n0.7  0.590  0.424\n0.5  0.560  0.501\n0.3  0.522  0.678\n========================================\n      R     P\nmean: 0.557 0.534\nstd:  0.028 0.106\n</code></pre>",
      "votes": 1,
      "replies": [
        {
          "id": 1658424,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-01-21T02:47:10.193000",
          "content": "<p>computation result of another 4 submission of different model:<br>\nWe have relatively small variance both on Recall and Precision.</p>\n<pre><code>         R      P\nrho\n0.7  0.614  0.554\n0.5  0.631  0.506\n0.3  0.625  0.521\n========================================\n      R     P\nmean: 0.623 0.527\nstd:  0.007 0.020\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1658520,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-01-21T05:14:27.773000",
          "content": "<p>Intuitive comprehension of why variance of precision is higher than that of recall is here.<br>\nSince F2 metrics focus more on recall than precision, our computation is more sensitive to recall than precision. Therefore we have higher variance in precision than recall.<br>\nOf course we need more math to prove this.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1661221,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-01-23T10:11:31.180000",
          "content": "<p>Solving the problem with least square fit, I have the following result:</p>\n<p>first model:</p>\n<pre><code>recall = 0.533\nprecision = 0.657\n</code></pre>\n<p>second model:</p>\n<pre><code>recall = 0.626\nprecision = 0.515\n</code></pre>\n<p>This time, we get more higher estimation of precision and slightly less recall in the first model.<br>\nFor the second model, estimated values are almost the same as for the former experiment.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1658384,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2022-01-21T01:28:55.097000",
      "content": "<h1>Simulation result</h1>\n<p>Assumptions:</p>\n<ul>\n<li>F2 score is rounded at decimal point 3</li>\n</ul>\n<p><a href=\"https://ibb.co/qCNVR9J\"><img src=\"https://i.ibb.co/xY1dMGF/Screen-Shot-2022-01-27-at-12-48-39.png\" alt=\"Screen-Shot-2022-01-27-at-12-48-39\"></a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1661207,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2022-01-23T10:04:53.730000",
      "content": "<h2>Tips: Estimate Recall and Precision with Least Square Method</h2>\n<p>If we have more than two observation of F2, we can solve this problem as least square regression.<br>\nGiven we obtain N equations (a) for different rho (rho_1, rho_2, …, rho_N), we have equations Ax = b where</p>\n<p>$$<br>\nA = \\begin{bmatrix}<br>\n4 &amp; \\rho_1 \\\\<br>\n4 &amp; \\rho_2 \\\\<br>\n… \\\\<br>\n4 &amp; \\rho_N \\\\<br>\n\\end{bmatrix},<br>\nb = \\begin{bmatrix}<br>\n\\rho_1 (5 / F_2^{\\rho_1} - 1) -4 \\\\<br>\n\\rho_2 (5 / F_2^{\\rho_2} - 1) -4 \\\\<br>\n… \\\\<br>\n\\rho_N (5 / F_2^{\\rho_N} - 1) -4 \\\\<br>\n\\end{bmatrix}, <br>\nx = \\begin{bmatrix}<br>\n\\frac{FN}{TP} \\\\<br>\n\\frac{FP}{TP} \\\\<br>\n\\end{bmatrix}<br>\n$$</p>\n<p>This is an over-determined system, so we can solve this problem by least-square regression[1].</p>\n<p>$$<br>\nmin_x\\|Ax - b\\|<br>\n$$</p>\n<p>The solution is</p>\n<p>$$<br>\nx = (A^\\top A)^{-1}A^\\top b<br>\n$$</p>\n<p>[1] <a href=\"https://en.wikipedia.org/wiki/Overdetermined_system#Approximate_solutions\" target=\"_blank\">https://en.wikipedia.org/wiki/Overdetermined_system#Approximate_solutions</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1658440,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2022-01-21T03:09:52.027000",
      "content": "<h2>Eg. estimating recall &amp; precision for the different train/test scales</h2>\n<p>I estimated recall and precision in the Leader board, and have result below.<br>\nIn this result, recall is higher when we infer larger scale than that in train, whereas precision is nearly the same.<br>\nOf course we need more sample because it depends on the model.<br>\nI don't confirm this result is generally applicable.</p>\n<p>train scale: x2.00, infer scale: x2.00</p>\n<pre><code>      R     P\nmean: 0.557 0.534\nstd:  0.028 0.106\n</code></pre>\n<p>train scale: x2.00, infer scale: x3.20 (x1.60 larger than train scale)</p>\n<pre><code>      R     P\nmean: 0.623 0.527\nstd:  0.007 0.020\n</code></pre>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1658827,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-01-21T10:44:31.367000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1658380": "# To be short\n\nWe can estimate recall and precision of your prediction model on public LB data.\nWhat we need is:\n\n1. Randomly dropping your model's prediction probability 1 - rho\n2. You need at least two submissions of different dropping probability (say, rho=1 and 0.5) and observe F2\n3. You can calculate precision and recall using the below code\n\n# Code \n\n```\ndef drop_pred(bboxes, confs, p_keep):\n    '''\n    randomly drop prediction for the probability (1 - p_keep).\n    '''\n    if p_keep == 1:\n        return bboxes, confs\n    bboxes = bboxes.copy()\n    confs = confs.copy()\n    assert len(bboxes) == len(confs)\n    pp = np.random.uniform(size=len(bboxes))\n    bboxes = bboxes[pp <= p_keep]\n    confs = confs[pp <= p_keep]\n    return bboxes, confs\n```\n\n```python\ndef estimate_recall_and_precision(f2, f2s, rhos):\n    recalls, precisions = [], []\n    for f2_rho, rho in zip(f2s, rhos):\n        fn_tp = 5 / 4 * rho / (1 - rho) * (1 / f2_rho - 1 / f2) - 1\n        fp_tp = 5 / (1 - rho) * (1 / f2 - 1 / f2_rho * rho) - 1\n        recall = 1 / (fn_tp + 1)\n        precision = 1 / (fp_tp + 1)\n        recalls.append(recall)\n        precisions.append(precision)\n\n    return recalls, precisions\n```\n\n# Detail\n\nIt was somehow tedious to verify, it might be some tiny mistake.\nBut I tested the final equation by simulation, so I think the code above is credible.\n\n---\n\nThe F2 score is calculated by equation below:\n\n$$\nF_2 = \\frac{5TP}{5TP+4FN+FP}\n$$\n\nIf we drop 1 - rho predictions, TP, FP, and FN changes to \n\n$$\n\\begin{eqnarray}\nTP^\\prime &=& \\rho TP \\\\\\\\\nFN^\\prime &=& FN + (1 - \\rho) TP \\\\\\\\\nFP^\\prime &=& \\rho FP\n\\end{eqnarray}\n$$\n\nAnd the new F2 score F2_rho is calculated as:\n\n$$\nF_2^\\rho = \\frac{5\\rho TP}{5\\rho TP+4(FN+(1-\\rho)TP)+\\rho FP}\n$$\n\nClearing out denominator and we get the equation of FN/TP and FP/TP:\n\n$$\n4 \\frac{FN}{TP} + \\rho \\frac{FP}{TP} = \\rho \\left( \\frac{5}{F_2^\\rho} - 1 \\right) - 4\n$$\n\nBecause the equation has two variables, we need at least two observation of F2_rho based on the different dropout ratio 1 - rho. Given we already have observation F2 (without dropout), we only require another observation of F_2_rho (e.g, rho=0.5).\n\nThen we have two linear equations:\n\n$$\n4 \\frac{FN}{TP} + \\frac{FP}{TP} = \\left( \\frac{5}{F_2} - 1 \\right) - 4 \\tag{a}\n$$\n$$\n4 \\frac{FN}{TP} + \\rho \\frac{FP}{TP} = \\rho \\left( \\frac{5}{F_2^\\rho} - 1 \\right) - 4 \\tag{b}\n$$\n\nSolving above equations, we get\n\n$$\n\\frac{FN}{TP}=\\frac{5}{4}\\frac{\\rho}{1-\\rho}\\left(\\frac{1}{F_2^{\\rho}}-\\frac{1}{F_2}\\right) - 1 \\tag{1}\n$$\n\n$$\n\\frac{FP}{TP}=\\frac{5}{1-\\rho}\\left( \\frac{1}{F_2} - \\frac{\\rho}{F_2^{\\rho}} \\right) -1 \\tag{2}\n$$\n\nAnd since recall=TP/(TP+FN) and precision=TP/(TP+FP), we can calculate recall and precision from this ratio.\n\n<a href=\"https://ibb.co/XpgvcGT\"><img src=\"https://i.ibb.co/3WqLgPH/Screen-Shot-2022-01-21-at-9-36-00.png\" alt=\"Screen-Shot-2022-01-21-at-9-36-00\" border=\"0\"></a>\n\n# Reference\n\nThis idea was inspired from below discussion by @hengck23:\nhttps://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300405#1656348\n\n----\n\n# Update Note\n\n* 2022/1/21 19:32 JST refactor equation & code: use more simpler form. Also do a simple degrade check by feeding few experiment data\n* 2022/1/22 18:19 JST simplify equation",
    "1658835": "make sure you have lots of background images (and many variations, e.g fish, bubble ...) to compute your model precision in validation.\n\nby comparing public test precision (or better still, comparing fp rate), we can guess the quality and number of background images.\n\ntop kagglers are good are modeling deep network and modeling unseen data. modeling unseen data is competition exploitation and is the key to winning. (expert kagglers don't experience shakeup)\n\nactually just private and public set is not enough for competition design to prevent exploitation.",
    "1658433": "this is human-based meta learning",
    "1658385": "I also make 4 submission of rho=(1.0, 0.7, 0.5, 0.3), and calculated the result.\nI observed high variance std=~0.1 on precision, but the variance of recall is relatively low (std=~0.03).\n\n```\n         R      P\nrho              \n0.7  0.590  0.424\n0.5  0.560  0.501\n0.3  0.522  0.678\n========================================\n      R     P\nmean: 0.557 0.534\nstd:  0.028 0.106\n```",
    "1658384": "# Simulation result\n\nAssumptions:\n* F2 score is rounded at decimal point 3\n\n<a href=\"https://ibb.co/qCNVR9J\"><img src=\"https://i.ibb.co/xY1dMGF/Screen-Shot-2022-01-27-at-12-48-39.png\" alt=\"Screen-Shot-2022-01-27-at-12-48-39\" border=\"0\"></a>",
    "1661207": "## Tips: Estimate Recall and Precision with Least Square Method\n\nIf we have more than two observation of F2, we can solve this problem as least square regression.\nGiven we obtain N equations (a) for different rho (rho_1, rho_2, ..., rho_N), we have equations Ax = b where\n\n$$\nA = \\begin{bmatrix}\n4 & \\rho_1 \\\\\\\\\n4 & \\rho_2 \\\\\\\\\n... \\\\\\\\\n4 & \\rho_N \\\\\\\\\n\\end{bmatrix},\nb = \\begin{bmatrix}\n\\rho_1 (5 / F_2^{\\rho_1} - 1) -4 \\\\\\\\\n\\rho_2 (5 / F_2^{\\rho_2} - 1) -4 \\\\\\\\\n... \\\\\\\\\n\\rho_N (5 / F_2^{\\rho_N} - 1) -4 \\\\\\\\\n\\end{bmatrix}, \nx = \\begin{bmatrix}\n\\frac{FN}{TP} \\\\\\\\\n\\frac{FP}{TP} \\\\\\\\\n\\end{bmatrix}\n$$\n\nThis is an over-determined system, so we can solve this problem by least-square regression[1].\n\n$$\nmin_x\\\\|Ax - b\\\\|\n$$\n\nThe solution is\n\n$$\nx = (A^\\top A)^{-1}A^\\top b\n$$\n\n[1] https://en.wikipedia.org/wiki/Overdetermined_system#Approximate_solutions",
    "1658440": "## Eg. estimating recall & precision for the different train/test scales\n\nI estimated recall and precision in the Leader board, and have result below.\nIn this result, recall is higher when we infer larger scale than that in train, whereas precision is nearly the same.\nOf course we need more sample because it depends on the model.\nI don't confirm this result is generally applicable.\n\ntrain scale: x2.00, infer scale: x2.00\n```\n      R     P\nmean: 0.557 0.534\nstd:  0.028 0.106\n```\n\ntrain scale: x2.00, infer scale: x3.20 (x1.60 larger than train scale)\n```\n      R     P\nmean: 0.623 0.527\nstd:  0.007 0.020\n```",
    "1658827": ""
  }
}