{
  "id": 322606,
  "title": "LB hack II: Methods for Estimating Statistics of Predictions to Public Test Data",
  "url": "/competitions/birdclef-2022/discussion/322606",
  "author_name": "Bilzard",
  "post_date": "2022-05-03T07:09:53.271000",
  "votes": 16,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I will share how to estimate the statistics (e.g. positive ratio) of a model's predictions on public test data.</p>\n<h2>When is it useful?</h2>\n<ul>\n<li>when you want to estimate positive/negative prediction ratio for the entire submission.</li>\n<li>when you want to estimate positive/negative prediction ratio for a particular species (or group of species).</li>\n</ul>\n<h2>Derivation</h2>\n<p>In the discussion that follows, we assume that the LB metrics are computed by balanced accuracy.</p>\n<p>The procedure is as follows.</p>\n<ol>\n<li>submit any model and observe the public LB value (we denote it m_0).</li>\n<li>in the second submission notebook, compute the statistics of the model's predictions (e.g., the positive prediction rate). Let this be p. Normalize p to the range [0, 1].</li>\n<li>in the same model as used in process 1, invert the positive predictions with random probability p.</li>\n<li>observe LB score of the second submission notebook (let this denote m_1).</li>\n</ol>\n<p>First, m_0 can be calculated from the definition of balanced accuracy as follows.<br>\n$$<br>\n2 m_0 = M_0 = \\text{TPR} + \\text{TNR} = \\text{TPR} + (1 - \\text{FPR}) = \\text{TPR} - \\text{FPR} + 1 \\tag{1}<br>\n$$</p>\n<p>And m_1 is computed as follows.</p>\n<p>$$<br>\n2 m_1 = M_1 = \\text{TPR}^\\prime - \\text{FPR}^\\prime + 1 = (1 - p) M_0 + p \\tag{2}<br>\n$$</p>\n<p>Note that TP, FP, TN, and FN in the second notebook are computed as (1-p)TP, (1-p)FP, TN + pFP, and FN + pTP, respectively (see attached figure).</p>\n<p>Solving equation (2) for p yields, </p>\n<p>$$<br>\np = \\frac{M_0 - M_1}{M_0 - 1} = \\frac{m_0 - m_1}{m_0 - 0.5} \\tag{3}<br>\n$$</p>\n<p>Substituting m_0 and m_1 into equation (3), we can estimate the original value of p.</p>\n<p><a href=\"https://ibb.co/BzxcQd5\"><img src=\"https://i.ibb.co/cgHkZMS/Screen-Shot-2022-05-03-at-15-49-03.png\" alt=\"Screen-Shot-2022-05-03-at-15-49-03\"></a></p>",
  "messages": [
    {
      "id": 1775578,
      "postDate": "2022-05-03T07:09:53.270Z",
      "content": "<p>I will share how to estimate the statistics (e.g. positive ratio) of a model's predictions on public test data.</p>\n<h2>When is it useful?</h2>\n<ul>\n<li>when you want to estimate positive/negative prediction ratio for the entire submission.</li>\n<li>when you want to estimate positive/negative prediction ratio for a particular species (or group of species).</li>\n</ul>\n<h2>Derivation</h2>\n<p>In the discussion that follows, we assume that the LB metrics are computed by balanced accuracy.</p>\n<p>The procedure is as follows.</p>\n<ol>\n<li>submit any model and observe the public LB value (we denote it m_0).</li>\n<li>in the second submission notebook, compute the statistics of the model's predictions (e.g., the positive prediction rate). Let this be p. Normalize p to the range [0, 1].</li>\n<li>in the same model as used in process 1, invert the positive predictions with random probability p.</li>\n<li>observe LB score of the second submission notebook (let this denote m_1).</li>\n</ol>\n<p>First, m_0 can be calculated from the definition of balanced accuracy as follows.<br>\n$$<br>\n2 m_0 = M_0 = \\text{TPR} + \\text{TNR} = \\text{TPR} + (1 - \\text{FPR}) = \\text{TPR} - \\text{FPR} + 1 \\tag{1}<br>\n$$</p>\n<p>And m_1 is computed as follows.</p>\n<p>$$<br>\n2 m_1 = M_1 = \\text{TPR}^\\prime - \\text{FPR}^\\prime + 1 = (1 - p) M_0 + p \\tag{2}<br>\n$$</p>\n<p>Note that TP, FP, TN, and FN in the second notebook are computed as (1-p)TP, (1-p)FP, TN + pFP, and FN + pTP, respectively (see attached figure).</p>\n<p>Solving equation (2) for p yields, </p>\n<p>$$<br>\np = \\frac{M_0 - M_1}{M_0 - 1} = \\frac{m_0 - m_1}{m_0 - 0.5} \\tag{3}<br>\n$$</p>\n<p>Substituting m_0 and m_1 into equation (3), we can estimate the original value of p.</p>\n<p><a href=\"https://ibb.co/BzxcQd5\"><img src=\"https://i.ibb.co/cgHkZMS/Screen-Shot-2022-05-03-at-15-49-03.png\" alt=\"Screen-Shot-2022-05-03-at-15-49-03\"></a></p>",
      "rawMarkdown": "I will share how to estimate the statistics (e.g. positive ratio) of a model's predictions on public test data.\n\n## When is it useful?\n\n* when you want to estimate positive/negative prediction ratio for the entire submission.\n* when you want to estimate positive/negative prediction ratio for a particular species (or group of species).\n\n## Derivation\n\nIn the discussion that follows, we assume that the LB metrics are computed by balanced accuracy.\n\nThe procedure is as follows.\n\n1. submit any model and observe the public LB value (we denote it m_0).\n2. in the second submission notebook, compute the statistics of the model's predictions (e.g., the positive prediction rate). Let this be p. Normalize p to the range [0, 1].\n3. in the same model as used in process 1, invert the positive predictions with random probability p.\n4. observe LB score of the second submission notebook (let this denote m_1).\n\nFirst, m_0 can be calculated from the definition of balanced accuracy as follows.\n$$\n2 m_0 = M_0 = \\text{TPR} + \\text{TNR} = \\text{TPR} + (1 - \\text{FPR}) = \\text{TPR} - \\text{FPR} + 1 \\tag{1}\n$$\n\nAnd m_1 is computed as follows.\n\n$$\n2 m_1 = M_1 = \\text{TPR}^\\prime - \\text{FPR}^\\prime + 1 = (1 - p) M_0 + p \\tag{2}\n$$\n\nNote that TP, FP, TN, and FN in the second notebook are computed as (1-p)TP, (1-p)FP, TN + pFP, and FN + pTP, respectively (see attached figure).\n\nSolving equation (2) for p yields, \n\n$$\np = \\frac{M_0 - M_1}{M_0 - 1} = \\frac{m_0 - m_1}{m_0 - 0.5} \\tag{3}\n$$\n\nSubstituting m_0 and m_1 into equation (3), we can estimate the original value of p.\n\n<a href=\"https://ibb.co/BzxcQd5\"><img src=\"https://i.ibb.co/cgHkZMS/Screen-Shot-2022-05-03-at-15-49-03.png\" alt=\"Screen-Shot-2022-05-03-at-15-49-03\" border=\"0\"></a>",
      "votes": 16
    },
    {
      "id": 1779032,
      "postDate": "2022-05-06T03:09:38.907Z",
      "content": "<p>I don't know why someone downvote this topic.<br>\nIt may certainly not be the method envisioned by the host, but I would think that if you find a hole in a competition like this, it would be better to share it so that all participants can use it rather than to let a few hackers monopolize it.<br>\nPlease understand that I am posting this in this context.</p>",
      "rawMarkdown": "I don't know why someone downvote this topic.\nIt may certainly not be the method envisioned by the host, but I would think that if you find a hole in a competition like this, it would be better to share it so that all participants can use it rather than to let a few hackers monopolize it.\nPlease understand that I am posting this in this context.",
      "votes": 1
    },
    {
      "id": 1778494,
      "postDate": "2022-05-05T11:17:19.750Z",
      "content": "<h2>How to gain more confirmation on hypothesis?</h2>\n<p>This technique has been used in other discussions.<br>\nTo guarantee the correctness of the hypothesis, I decided to increase the sample of estimates for the known p.</p>\n<table>\n<thead>\n<tr>\n<th>p_true</th>\n<th>m_0</th>\n<th>m_1</th>\n<th>p_est</th>\n<th>error</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.15</td>\n<td>0.71</td>\n<td>0.68</td>\n<td>0.143</td>\n<td>0.007</td>\n</tr>\n<tr>\n<td>0.4</td>\n<td>0.71</td>\n<td>0.63</td>\n<td>0.381</td>\n<td>0.019</td>\n</tr>\n<tr>\n<td>0.6</td>\n<td>0.71</td>\n<td>0.57</td>\n<td>0.667</td>\n<td>0.067</td>\n</tr>\n<tr>\n<td>0.8</td>\n<td>0.71</td>\n<td>0.54</td>\n<td>0.810</td>\n<td>0.010</td>\n</tr>\n<tr>\n<td>0.905</td>\n<td>0.71</td>\n<td>0.50</td>\n<td>1.000</td>\n<td>0.095</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "## How to gain more confirmation on hypothesis?\n\nThis technique has been used in other discussions.\nTo guarantee the correctness of the hypothesis, I decided to increase the sample of estimates for the known p.\n\np_true|m_0|m_1|p_est|error\n--|--|--|--|--\n0.15|0.71|0.68|0.143|0.007\n0.4|0.71|0.63|0.381|0.019\n0.6|0.71|0.57|0.667|0.067\n0.8|0.71|0.54|0.810|0.010\n0.905|0.71|0.50|1.000|0.095",
      "votes": 1
    },
    {
      "id": 1775994,
      "postDate": "2022-05-03T15:38:54.433Z",
      "content": "<h2>Experiment: estimation of positive example prediction rates for a public notebook</h2>\n<ol>\n<li>split the scored species into four subgroups: <em>top5</em>, <em>mid_top5</em>, <em>mid_low5</em>, <em>low6</em> in descending order of sample size</li>\n<li>measured the positive prediction ratio p=(TP + FP) / (TP + FP + TN + FN) for each group</li>\n<li>from the notebook[1] scores m_0, m_1, estimate p using this discussion's method</li>\n</ol>\n<h2>Sample code</h2>\n<pre><code>def p(m_0, m_1):\n    return (m_0 - m_1) / (m_0 - 0.5)\n</code></pre>\n<h2>Results</h2>\n<p>The positive prediction ratio of <em>top5</em> species was <strong>0.76</strong> and that of <em>low6</em> was <strong>0.14</strong>.<br>\nAs stated in the discussion[2], the positive prediction ratio of <em>top5</em> species are much higher than that of <em>low6</em>.</p>\n<pre><code>top5, mid_top5, mid_low5, low6 =\n(['skylar', 'houfin', 'jabwar', 'warwhe1', 'yefcan'],\n ['apapan', 'iiwi', 'omao', 'hawama', 'hawcre'],\n ['barpet', 'akiapo', 'elepai', 'aniani', 'hawgoo'],\n ['ercfra', 'hawpet1', 'puaioh', 'hawhaw', 'crehon', 'maupar'])\n\nm_0: 0.71\nm_1 for top5: 0.55\nm_1 for mid_top5: 0.57\nm_1 for mid_low5: 0.65\nm_1 for low6: 0.68\n\nestimated p of top5: 0.762\nestimated p of mid_top5: 0.667\nestimated p of mid_low5: 0.286\nestimated p of low6: 0.143\n</code></pre>\n<h2>Reference</h2>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/code/kaerunantoka/birdclef2022-ex005-f0-infer\" target=\"_blank\">https://www.kaggle.com/code/kaerunantoka/birdclef2022-ex005-f0-infer</a></li>\n<li>[2] <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/322419#1774366\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/322419#1774366</a></li>\n</ul>",
      "rawMarkdown": "## Experiment: estimation of positive example prediction rates for a public notebook\n\n1. split the scored species into four subgroups: *top5*, *mid_top5*, *mid_low5*, *low6* in descending order of sample size\n2. measured the positive prediction ratio p=(TP + FP) / (TP + FP + TN + FN) for each group\n3. from the notebook[1] scores m_0, m_1, estimate p using this discussion's method\n\n## Sample code\n\n```\ndef p(m_0, m_1):\n    return (m_0 - m_1) / (m_0 - 0.5)\n```\n\n## Results\n\nThe positive prediction ratio of *top5* species was **0.76** and that of *low6* was **0.14**.\nAs stated in the discussion[2], the positive prediction ratio of *top5* species are much higher than that of *low6*.\n\n```\ntop5, mid_top5, mid_low5, low6 =\n(['skylar', 'houfin', 'jabwar', 'warwhe1', 'yefcan'],\n ['apapan', 'iiwi', 'omao', 'hawama', 'hawcre'],\n ['barpet', 'akiapo', 'elepai', 'aniani', 'hawgoo'],\n ['ercfra', 'hawpet1', 'puaioh', 'hawhaw', 'crehon', 'maupar'])\n\nm_0: 0.71\nm_1 for top5: 0.55\nm_1 for mid_top5: 0.57\nm_1 for mid_low5: 0.65\nm_1 for low6: 0.68\n\nestimated p of top5: 0.762\nestimated p of mid_top5: 0.667\nestimated p of mid_low5: 0.286\nestimated p of low6: 0.143\n```\n\n## Reference\n\n- [1] https://www.kaggle.com/code/kaerunantoka/birdclef2022-ex005-f0-infer\n- [2] https://www.kaggle.com/competitions/birdclef-2022/discussion/322419#1774366\n",
      "votes": 1
    },
    {
      "id": 1775743,
      "postDate": "2022-05-03T10:53:32.540Z",
      "content": "<h2>Note: Theoretical error of estimate of p</h2>\n<p>The theoretical error of the estimate of p can be calculated by the following equation.</p>\n<p>$$<br>\np_\\text{error} = \\frac{0.01}{m_0 - 0.5} = 0.0476, (\\text{where } m_0=0.71)<br>\n$$</p>\n<p>Note that the above is a calculation of the upper bound of the error, not the expected value.</p>",
      "rawMarkdown": "## Note: Theoretical error of estimate of p\n\nThe theoretical error of the estimate of p can be calculated by the following equation.\n\n$$\np_\\text{error} = \\frac{0.01}{m_0 - 0.5} = 0.0476, (\\text{where } m_0=0.71)\n$$\n\nNote that the above is a calculation of the upper bound of the error, not the expected value.",
      "votes": 1
    },
    {
      "id": 1775738,
      "postDate": "2022-05-03T10:41:36Z",
      "content": "<h2>Experiment: Testing the hypothesis</h2>\n<p>To test the above hypothesis, the following experiments were conducted:</p>\n<ul>\n<li>Fixed at p=0.4 and observe m_0 and m_1. Calculate the estimated value of p calculated from equation (3) and check how close it is to the original value.</li>\n</ul>\n<h2>Result</h2>\n<pre><code>m_0 = 0.71\nm_1 = 0.63\n\np_est = 0.381\np_error/p_true = 0.0475 (4.75%)\n</code></pre>\n<h2>Conclusion</h2>\n<p>The estimates in the experiment were confirmed to be correct within 5% precision. This fact proves to some extent the correctness of the hypotheses in this discussion.<br>\n(To increase the reliability of the estimates, we need to increase the number of measurements.)</p>\n<h2>Code used for above verification</h2>\n<pre><code>def random_invert_pos_target(input_df, p=1.0):\n    assert (p &gt;= 0.0) and (p &lt;= 1.0), p\n    tmp_df = input_df.copy()\n    pos_df = tmp_df.query(\"target == True\").reset_index()\n    n_rows = len(pos_df)\n    n_inverted = int(n_rows * p)\n    idxs = np.random.permutation(n_rows)[:n_inverted]\n    pos_df.loc[idxs, \"target\"] = pos_df.loc[idxs, \"target\"].apply(lambda x: not (x))\n    pos_df = pos_df.set_index(\"index\")\n    pos_idxs = pos_df.index\n    tmp_df.loc[pos_idxs, \"target\"] = pos_df.loc[pos_idxs, \"target\"]\n\n    return tmp_df\n</code></pre>\n<pre><code>sample_submission = ...  # your model's prediction\np = 0.4\ntmp_df = random_invert_pos_target(sample_submission, p=p)\nsample_submission[\"target\"] = tmp_df[\"target\"]\n</code></pre>",
      "rawMarkdown": "## Experiment: Testing the hypothesis\n\nTo test the above hypothesis, the following experiments were conducted:\n* Fixed at p=0.4 and observe m_0 and m_1. Calculate the estimated value of p calculated from equation (3) and check how close it is to the original value.\n\n## Result\n\n```\nm_0 = 0.71\nm_1 = 0.63\n\np_est = 0.381\np_error/p_true = 0.0475 (4.75%)\n```\n\n## Conclusion\n\nThe estimates in the experiment were confirmed to be correct within 5% precision. This fact proves to some extent the correctness of the hypotheses in this discussion.\n(To increase the reliability of the estimates, we need to increase the number of measurements.)\n\n## Code used for above verification\n\n```\ndef random_invert_pos_target(input_df, p=1.0):\n    assert (p >= 0.0) and (p <= 1.0), p\n    tmp_df = input_df.copy()\n    pos_df = tmp_df.query(\"target == True\").reset_index()\n    n_rows = len(pos_df)\n    n_inverted = int(n_rows * p)\n    idxs = np.random.permutation(n_rows)[:n_inverted]\n    pos_df.loc[idxs, \"target\"] = pos_df.loc[idxs, \"target\"].apply(lambda x: not (x))\n    pos_df = pos_df.set_index(\"index\")\n    pos_idxs = pos_df.index\n    tmp_df.loc[pos_idxs, \"target\"] = pos_df.loc[pos_idxs, \"target\"]\n\n    return tmp_df\n```\n\n```\nsample_submission = ...  # your model's prediction\np = 0.4\ntmp_df = random_invert_pos_target(sample_submission, p=p)\nsample_submission[\"target\"] = tmp_df[\"target\"]\n```",
      "votes": 1
    },
    {
      "id": 1775588,
      "postDate": "2022-05-03T07:28:04.593Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1779032,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2022-05-06T03:09:38.907000",
      "content": "<p>I don't know why someone downvote this topic.<br>\nIt may certainly not be the method envisioned by the host, but I would think that if you find a hole in a competition like this, it would be better to share it so that all participants can use it rather than to let a few hackers monopolize it.<br>\nPlease understand that I am posting this in this context.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1778494,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2022-05-05T11:17:19.750000",
      "content": "<h2>How to gain more confirmation on hypothesis?</h2>\n<p>This technique has been used in other discussions.<br>\nTo guarantee the correctness of the hypothesis, I decided to increase the sample of estimates for the known p.</p>\n<table>\n<thead>\n<tr>\n<th>p_true</th>\n<th>m_0</th>\n<th>m_1</th>\n<th>p_est</th>\n<th>error</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.15</td>\n<td>0.71</td>\n<td>0.68</td>\n<td>0.143</td>\n<td>0.007</td>\n</tr>\n<tr>\n<td>0.4</td>\n<td>0.71</td>\n<td>0.63</td>\n<td>0.381</td>\n<td>0.019</td>\n</tr>\n<tr>\n<td>0.6</td>\n<td>0.71</td>\n<td>0.57</td>\n<td>0.667</td>\n<td>0.067</td>\n</tr>\n<tr>\n<td>0.8</td>\n<td>0.71</td>\n<td>0.54</td>\n<td>0.810</td>\n<td>0.010</td>\n</tr>\n<tr>\n<td>0.905</td>\n<td>0.71</td>\n<td>0.50</td>\n<td>1.000</td>\n<td>0.095</td>\n</tr>\n</tbody>\n</table>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1775994,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2022-05-03T15:38:54.433000",
      "content": "<h2>Experiment: estimation of positive example prediction rates for a public notebook</h2>\n<ol>\n<li>split the scored species into four subgroups: <em>top5</em>, <em>mid_top5</em>, <em>mid_low5</em>, <em>low6</em> in descending order of sample size</li>\n<li>measured the positive prediction ratio p=(TP + FP) / (TP + FP + TN + FN) for each group</li>\n<li>from the notebook[1] scores m_0, m_1, estimate p using this discussion's method</li>\n</ol>\n<h2>Sample code</h2>\n<pre><code>def p(m_0, m_1):\n    return (m_0 - m_1) / (m_0 - 0.5)\n</code></pre>\n<h2>Results</h2>\n<p>The positive prediction ratio of <em>top5</em> species was <strong>0.76</strong> and that of <em>low6</em> was <strong>0.14</strong>.<br>\nAs stated in the discussion[2], the positive prediction ratio of <em>top5</em> species are much higher than that of <em>low6</em>.</p>\n<pre><code>top5, mid_top5, mid_low5, low6 =\n(['skylar', 'houfin', 'jabwar', 'warwhe1', 'yefcan'],\n ['apapan', 'iiwi', 'omao', 'hawama', 'hawcre'],\n ['barpet', 'akiapo', 'elepai', 'aniani', 'hawgoo'],\n ['ercfra', 'hawpet1', 'puaioh', 'hawhaw', 'crehon', 'maupar'])\n\nm_0: 0.71\nm_1 for top5: 0.55\nm_1 for mid_top5: 0.57\nm_1 for mid_low5: 0.65\nm_1 for low6: 0.68\n\nestimated p of top5: 0.762\nestimated p of mid_top5: 0.667\nestimated p of mid_low5: 0.286\nestimated p of low6: 0.143\n</code></pre>\n<h2>Reference</h2>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/code/kaerunantoka/birdclef2022-ex005-f0-infer\" target=\"_blank\">https://www.kaggle.com/code/kaerunantoka/birdclef2022-ex005-f0-infer</a></li>\n<li>[2] <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/322419#1774366\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/322419#1774366</a></li>\n</ul>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1775743,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2022-05-03T10:53:32.540000",
      "content": "<h2>Note: Theoretical error of estimate of p</h2>\n<p>The theoretical error of the estimate of p can be calculated by the following equation.</p>\n<p>$$<br>\np_\\text{error} = \\frac{0.01}{m_0 - 0.5} = 0.0476, (\\text{where } m_0=0.71)<br>\n$$</p>\n<p>Note that the above is a calculation of the upper bound of the error, not the expected value.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1775738,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2022-05-03T10:41:36",
      "content": "<h2>Experiment: Testing the hypothesis</h2>\n<p>To test the above hypothesis, the following experiments were conducted:</p>\n<ul>\n<li>Fixed at p=0.4 and observe m_0 and m_1. Calculate the estimated value of p calculated from equation (3) and check how close it is to the original value.</li>\n</ul>\n<h2>Result</h2>\n<pre><code>m_0 = 0.71\nm_1 = 0.63\n\np_est = 0.381\np_error/p_true = 0.0475 (4.75%)\n</code></pre>\n<h2>Conclusion</h2>\n<p>The estimates in the experiment were confirmed to be correct within 5% precision. This fact proves to some extent the correctness of the hypotheses in this discussion.<br>\n(To increase the reliability of the estimates, we need to increase the number of measurements.)</p>\n<h2>Code used for above verification</h2>\n<pre><code>def random_invert_pos_target(input_df, p=1.0):\n    assert (p &gt;= 0.0) and (p &lt;= 1.0), p\n    tmp_df = input_df.copy()\n    pos_df = tmp_df.query(\"target == True\").reset_index()\n    n_rows = len(pos_df)\n    n_inverted = int(n_rows * p)\n    idxs = np.random.permutation(n_rows)[:n_inverted]\n    pos_df.loc[idxs, \"target\"] = pos_df.loc[idxs, \"target\"].apply(lambda x: not (x))\n    pos_df = pos_df.set_index(\"index\")\n    pos_idxs = pos_df.index\n    tmp_df.loc[pos_idxs, \"target\"] = pos_df.loc[pos_idxs, \"target\"]\n\n    return tmp_df\n</code></pre>\n<pre><code>sample_submission = ...  # your model's prediction\np = 0.4\ntmp_df = random_invert_pos_target(sample_submission, p=p)\nsample_submission[\"target\"] = tmp_df[\"target\"]\n</code></pre>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1775588,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-03T07:28:04.593000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1775578": "I will share how to estimate the statistics (e.g. positive ratio) of a model's predictions on public test data.\n\n## When is it useful?\n\n* when you want to estimate positive/negative prediction ratio for the entire submission.\n* when you want to estimate positive/negative prediction ratio for a particular species (or group of species).\n\n## Derivation\n\nIn the discussion that follows, we assume that the LB metrics are computed by balanced accuracy.\n\nThe procedure is as follows.\n\n1. submit any model and observe the public LB value (we denote it m_0).\n2. in the second submission notebook, compute the statistics of the model's predictions (e.g., the positive prediction rate). Let this be p. Normalize p to the range [0, 1].\n3. in the same model as used in process 1, invert the positive predictions with random probability p.\n4. observe LB score of the second submission notebook (let this denote m_1).\n\nFirst, m_0 can be calculated from the definition of balanced accuracy as follows.\n$$\n2 m_0 = M_0 = \\text{TPR} + \\text{TNR} = \\text{TPR} + (1 - \\text{FPR}) = \\text{TPR} - \\text{FPR} + 1 \\tag{1}\n$$\n\nAnd m_1 is computed as follows.\n\n$$\n2 m_1 = M_1 = \\text{TPR}^\\prime - \\text{FPR}^\\prime + 1 = (1 - p) M_0 + p \\tag{2}\n$$\n\nNote that TP, FP, TN, and FN in the second notebook are computed as (1-p)TP, (1-p)FP, TN + pFP, and FN + pTP, respectively (see attached figure).\n\nSolving equation (2) for p yields, \n\n$$\np = \\frac{M_0 - M_1}{M_0 - 1} = \\frac{m_0 - m_1}{m_0 - 0.5} \\tag{3}\n$$\n\nSubstituting m_0 and m_1 into equation (3), we can estimate the original value of p.\n\n<a href=\"https://ibb.co/BzxcQd5\"><img src=\"https://i.ibb.co/cgHkZMS/Screen-Shot-2022-05-03-at-15-49-03.png\" alt=\"Screen-Shot-2022-05-03-at-15-49-03\" border=\"0\"></a>",
    "1779032": "I don't know why someone downvote this topic.\nIt may certainly not be the method envisioned by the host, but I would think that if you find a hole in a competition like this, it would be better to share it so that all participants can use it rather than to let a few hackers monopolize it.\nPlease understand that I am posting this in this context.",
    "1778494": "## How to gain more confirmation on hypothesis?\n\nThis technique has been used in other discussions.\nTo guarantee the correctness of the hypothesis, I decided to increase the sample of estimates for the known p.\n\np_true|m_0|m_1|p_est|error\n--|--|--|--|--\n0.15|0.71|0.68|0.143|0.007\n0.4|0.71|0.63|0.381|0.019\n0.6|0.71|0.57|0.667|0.067\n0.8|0.71|0.54|0.810|0.010\n0.905|0.71|0.50|1.000|0.095",
    "1775994": "## Experiment: estimation of positive example prediction rates for a public notebook\n\n1. split the scored species into four subgroups: *top5*, *mid_top5*, *mid_low5*, *low6* in descending order of sample size\n2. measured the positive prediction ratio p=(TP + FP) / (TP + FP + TN + FN) for each group\n3. from the notebook[1] scores m_0, m_1, estimate p using this discussion's method\n\n## Sample code\n\n```\ndef p(m_0, m_1):\n    return (m_0 - m_1) / (m_0 - 0.5)\n```\n\n## Results\n\nThe positive prediction ratio of *top5* species was **0.76** and that of *low6* was **0.14**.\nAs stated in the discussion[2], the positive prediction ratio of *top5* species are much higher than that of *low6*.\n\n```\ntop5, mid_top5, mid_low5, low6 =\n(['skylar', 'houfin', 'jabwar', 'warwhe1', 'yefcan'],\n ['apapan', 'iiwi', 'omao', 'hawama', 'hawcre'],\n ['barpet', 'akiapo', 'elepai', 'aniani', 'hawgoo'],\n ['ercfra', 'hawpet1', 'puaioh', 'hawhaw', 'crehon', 'maupar'])\n\nm_0: 0.71\nm_1 for top5: 0.55\nm_1 for mid_top5: 0.57\nm_1 for mid_low5: 0.65\nm_1 for low6: 0.68\n\nestimated p of top5: 0.762\nestimated p of mid_top5: 0.667\nestimated p of mid_low5: 0.286\nestimated p of low6: 0.143\n```\n\n## Reference\n\n- [1] https://www.kaggle.com/code/kaerunantoka/birdclef2022-ex005-f0-infer\n- [2] https://www.kaggle.com/competitions/birdclef-2022/discussion/322419#1774366\n",
    "1775743": "## Note: Theoretical error of estimate of p\n\nThe theoretical error of the estimate of p can be calculated by the following equation.\n\n$$\np_\\text{error} = \\frac{0.01}{m_0 - 0.5} = 0.0476, (\\text{where } m_0=0.71)\n$$\n\nNote that the above is a calculation of the upper bound of the error, not the expected value.",
    "1775738": "## Experiment: Testing the hypothesis\n\nTo test the above hypothesis, the following experiments were conducted:\n* Fixed at p=0.4 and observe m_0 and m_1. Calculate the estimated value of p calculated from equation (3) and check how close it is to the original value.\n\n## Result\n\n```\nm_0 = 0.71\nm_1 = 0.63\n\np_est = 0.381\np_error/p_true = 0.0475 (4.75%)\n```\n\n## Conclusion\n\nThe estimates in the experiment were confirmed to be correct within 5% precision. This fact proves to some extent the correctness of the hypotheses in this discussion.\n(To increase the reliability of the estimates, we need to increase the number of measurements.)\n\n## Code used for above verification\n\n```\ndef random_invert_pos_target(input_df, p=1.0):\n    assert (p >= 0.0) and (p <= 1.0), p\n    tmp_df = input_df.copy()\n    pos_df = tmp_df.query(\"target == True\").reset_index()\n    n_rows = len(pos_df)\n    n_inverted = int(n_rows * p)\n    idxs = np.random.permutation(n_rows)[:n_inverted]\n    pos_df.loc[idxs, \"target\"] = pos_df.loc[idxs, \"target\"].apply(lambda x: not (x))\n    pos_df = pos_df.set_index(\"index\")\n    pos_idxs = pos_df.index\n    tmp_df.loc[pos_idxs, \"target\"] = pos_df.loc[pos_idxs, \"target\"]\n\n    return tmp_df\n```\n\n```\nsample_submission = ...  # your model's prediction\np = 0.4\ntmp_df = random_invert_pos_target(sample_submission, p=p)\nsample_submission[\"target\"] = tmp_df[\"target\"]\n```",
    "1775588": ""
  }
}