{
  "id": 441139,
  "title": "Trouble reproducing evaluation metric (MRRMSE) locally?",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/441139",
  "author_name": "qihuaz",
  "post_date": "2023-09-17T17:35:00.689000",
  "votes": 5,
  "comment_count": 7,
  "views": 0,
  "content": "<p>hello all, I am tring to reproduce the evaluation metric Mean Rowwise Root Mean Squared Error (MRRMSE) with the train data. But the result looks much different from the public leaderboard score.</p>\n<p>Just as a sanity check, I try to see what the score will be if I predict all 0 for B cells and Myeloid cells, evaluated by the availble train data (excluding positive controls). I am expecting to get something around 0.666 (what you would get by submitting all 0s to the public leaderboard).</p>\n<p>Since train is different than test, so I am expecting the value around but not exactly 0.666, maybe in the 0.3~0.9 range. Butinstad I got 2.0345?</p>\n<p><strong>Am I doing something wrong? Or the public test data is just dramastically different from the train data?</strong></p>\n<p>Here's how I got the 2.0345 from the processed train data:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1716193%2F5e6508336e1af45bb6b04fed4404db2d%2FScreenshot%202023-09-17%20133355.png?generation=1694972045873128&amp;alt=media\" alt=\"\"><br>\nthen</p>\n<p>$$\\textrm{MRRMSE} = \\frac{1}{R}\\sum_{i=1}^R\\left(\\frac{1}{n} \\sum_{j=1}^{n} (y_{ij} - \\widehat{y}_{ij})^2\\right)^{1/2}$$</p>\n<pre><code>(\n    (filtered_df_train.iloc[:, :] **      \n    ).mean(axis=) **                    \n).mean()                                    \n</code></pre>\n<p>I hope I am wrong here.. otherwise it would just a lot of guessing game rather than analysis if train set and the test set are distributed so differently… </p>\n<p>The notebook:<br>\n<a href=\"https://www.kaggle.com/code/qihuaz/mean-rowwise-root-mean-squared-error?scriptVersionId=143324165\" target=\"_blank\">https://www.kaggle.com/code/qihuaz/mean-rowwise-root-mean-squared-error?scriptVersionId=143324165</a></p>",
  "messages": [
    {
      "id": 2443398,
      "postDate": "2023-09-17T17:35:00.690Z",
      "content": "<p>hello all, I am tring to reproduce the evaluation metric Mean Rowwise Root Mean Squared Error (MRRMSE) with the train data. But the result looks much different from the public leaderboard score.</p>\n<p>Just as a sanity check, I try to see what the score will be if I predict all 0 for B cells and Myeloid cells, evaluated by the availble train data (excluding positive controls). I am expecting to get something around 0.666 (what you would get by submitting all 0s to the public leaderboard).</p>\n<p>Since train is different than test, so I am expecting the value around but not exactly 0.666, maybe in the 0.3~0.9 range. Butinstad I got 2.0345?</p>\n<p><strong>Am I doing something wrong? Or the public test data is just dramastically different from the train data?</strong></p>\n<p>Here's how I got the 2.0345 from the processed train data:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1716193%2F5e6508336e1af45bb6b04fed4404db2d%2FScreenshot%202023-09-17%20133355.png?generation=1694972045873128&amp;alt=media\" alt=\"\"><br>\nthen</p>\n<p>$$\\textrm{MRRMSE} = \\frac{1}{R}\\sum_{i=1}^R\\left(\\frac{1}{n} \\sum_{j=1}^{n} (y_{ij} - \\widehat{y}_{ij})^2\\right)^{1/2}$$</p>\n<pre><code>(\n    (filtered_df_train.iloc[:, :] **      \n    ).mean(axis=) **                    \n).mean()                                    \n</code></pre>\n<p>I hope I am wrong here.. otherwise it would just a lot of guessing game rather than analysis if train set and the test set are distributed so differently… </p>\n<p>The notebook:<br>\n<a href=\"https://www.kaggle.com/code/qihuaz/mean-rowwise-root-mean-squared-error?scriptVersionId=143324165\" target=\"_blank\">https://www.kaggle.com/code/qihuaz/mean-rowwise-root-mean-squared-error?scriptVersionId=143324165</a></p>",
      "rawMarkdown": "hello all, I am tring to reproduce the evaluation metric Mean Rowwise Root Mean Squared Error (MRRMSE) with the train data. But the result looks much different from the public leaderboard score.\n\nJust as a sanity check, I try to see what the score will be if I predict all 0 for B cells and Myeloid cells, evaluated by the availble train data (excluding positive controls). I am expecting to get something around 0.666 (what you would get by submitting all 0s to the public leaderboard).\n\nSince train is different than test, so I am expecting the value around but not exactly 0.666, maybe in the 0.3~0.9 range. Butinstad I got 2.0345?\n\n**Am I doing something wrong? Or the public test data is just dramastically different from the train data?**\n\nHere's how I got the 2.0345 from the processed train data:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1716193%2F5e6508336e1af45bb6b04fed4404db2d%2FScreenshot%202023-09-17%20133355.png?generation=1694972045873128&alt=media)\nthen\n\n$$\\textrm{MRRMSE} = \\frac{1}{R}\\sum_{i=1}^R\\left(\\frac{1}{n} \\sum_{j=1}^{n} (y_{ij} - \\widehat{y}_{ij})^2\\right)^{1/2}$$\n\n```python\n(\n    (filtered_df_train.iloc[:, 5:] ** 2     # (y_ij - yhat_{ij}) ^ 2\n    ).mean(axis=1) ** 0.5                   # take the mean of n genes then sqaure root\n).mean()                                    # take the mean by R\n```\n\nI hope I am wrong here.. otherwise it would just a lot of guessing game rather than analysis if train set and the test set are distributed so differently... \n\nThe notebook:\nhttps://www.kaggle.com/code/qihuaz/mean-rowwise-root-mean-squared-error?scriptVersionId=143324165",
      "votes": 5
    },
    {
      "id": 2445109,
      "postDate": "2023-09-18T16:14:29.450Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/qihuaz\" target=\"_blank\">@qihuaz</a>!</p>\n<p>The following is the naive implementation of MRRMSE:</p>\n<pre><code> () -&gt; np.float64:\n    \n     Y_true.shape == Y_pred.shape\n    R, n = Y_true.shape\n    result = \n     i  (R):\n        row = \n         j  (n):\n            row += (Y_true[i, j] - Y_pred[i, j]) ** \n        row = np.sqrt(row / n)\n        result += row\n     result / R\n</code></pre>\n<p>But as expected it is slow because of Python's for loop:</p>\n<pre><code>(df_train[target_cols].values, df_train[target_cols].values)\n times: user . s, sys: . ms, total: . s\n time: . s\n</code></pre>\n<p>The vectorized version looks like:</p>\n<pre><code> () -&gt; np.float64:\n    \n     Y_true.shape == Y_pred.shape\n    R, _ = Y_true.shape\n    result = \n     i  (R):\n        row_sum = np.square(Y_true[i, :] - Y_pred[i, :]).mean()\n        result += row_sum\n     result / R\n</code></pre>\n<pre><code>(df_train[target_cols].values, df_train[target_cols].values)\n times: user . ms, sys: . ms, total: . ms\n time: . ms\n</code></pre>\n<p>I hope I understood your issue properly and snippets from above may be of use.</p>",
      "rawMarkdown": "Hi @qihuaz!\n\nThe following is the naive implementation of MRRMSE:\n```Python\ndef mrrmse(Y_true:np.ndarray, Y_pred:np.ndarray) -> np.float64:\n    \"\"\"Naive implementation of MRRMSE.\"\"\"\n    assert Y_true.shape == Y_pred.shape\n    R, n = Y_true.shape\n    result = 0\n    for i in range(R):\n        row = 0\n        for j in range(n):\n            row += (Y_true[i, j] - Y_pred[i, j]) ** 2\n        row = np.sqrt(row / n)\n        result += row\n    return result / R\n```\nBut as expected it is slow because of Python's for loop:\n```%%time\nmrrmse(df_train[target_cols].values, df_train[target_cols].values)\nCPU times: user 4.31 s, sys: 4.72 ms, total: 4.32 s\nWall time: 4.32 s\n```\n\nThe vectorized version looks like:\n```Python\ndef mrrmse_vectorized(Y_true:np.ndarray, Y_pred:np.ndarray) -> np.float64:\n    \"\"\"Vectorized implementation of MRRMSE.\"\"\"\n    assert Y_true.shape == Y_pred.shape\n    R, _ = Y_true.shape\n    result = 0\n    for i in range(R):\n        row_sum = np.square(Y_true[i, :] - Y_pred[i, :]).mean()\n        result += row_sum\n    return result / R\n```\n\n```%%time\nmrrmse_vectorized(df_train[target_cols].values, df_train[target_cols].values)\nCPU times: user 67.8 ms, sys: 15.9 ms, total: 83.8 ms\nWall time: 82.6 ms\n```\n\n\nI hope I understood your issue properly and snippets from above may be of use.",
      "votes": 1,
      "replies": [
        {
          "id": 2445344,
          "postDate": "2023-09-18T18:50:14.900Z",
          "content": "<p>Thanks for the implementation. </p>\n<p>My code was trying to do a quick and dirty calculation of the mrrmse result if 0s are submitted for all prediction. I didn't run yours, but they should come to the same number. </p>\n<p>My issue is that it seems like under this metric, the result is not comparable accross train/public test/private test (please see my reply to <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> for more details).</p>",
          "rawMarkdown": "Thanks for the implementation. \n\nMy code was trying to do a quick and dirty calculation of the mrrmse result if 0s are submitted for all prediction. I didn't run yours, but they should come to the same number. \n\nMy issue is that it seems like under this metric, the result is not comparable accross train/public test/private test (please see my reply to @danielburkhardt for more details).",
          "replies": [
            {
              "id": 2445676,
              "postDate": "2023-09-19T02:57:15.873Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        },
        {
          "id": 2445678,
          "postDate": "2023-09-19T02:57:55.630Z",
          "content": "<p><a href=\"https://www.kaggle.com/calmscout\" target=\"_blank\">@calmscout</a> I think the <code>np.sqrt</code> are missing in the <code>mrrmse_vectorized</code>?</p>",
          "rawMarkdown": "@calmscout I think the `np.sqrt` are missing in the `mrrmse_vectorized`?",
          "votes": 1
        },
        {
          "id": 2454047,
          "postDate": "2023-09-24T14:25:01.507Z",
          "content": "<p>Here is the code more efficent and fixing the missing root:</p>\n<pre><code> () -&gt; np.float64:\n    squared_diffs = np.square(Y_true - Y_pred)\n    row_means = squared_diffs.mean(axis=)\n    root_means = np.sqrt(row_means)\n     root_means.mean()\n</code></pre>",
          "rawMarkdown": "Here is the code more efficent and fixing the missing root:\n\n```python\ndef rrmse(Y_true: np.ndarray, Y_pred: np.ndarray) -> np.float64:\n    squared_diffs = np.square(Y_true - Y_pred)\n    row_means = squared_diffs.mean(axis=1)\n    root_means = np.sqrt(row_means)\n    return root_means.mean()\n```",
          "votes": 2
        }
      ]
    },
    {
      "id": 2444970,
      "postDate": "2023-09-18T15:02:28.303Z",
      "content": "<p>Perhaps <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> can help?</p>",
      "rawMarkdown": "Perhaps @ryanholbrook can help?",
      "replies": [
        {
          "id": 2445355,
          "postDate": "2023-09-18T18:59:49.860Z",
          "content": "<blockquote>\n  <p>Perhaps <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> can help?</p>\n</blockquote>\n<p>Upon a bit of further investigation. It seems like the in the train data, a compound <code>MLN 2238</code> that causes high DEs can \"single-handedly\" contirbutes to a bulk part of the MRRMSE.</p>\n<p>Since MRRMSE simply average over rows (instead of somehow limitting the influence of a single row/compound), a single outlier compound like <code>MLN 2238</code> can outweight all other compounds combined in its contribution to the MRRMSE in the train data. <strong>Judging from the 0.666 MRRMSE score for the sample submission (all 0), there is no similar such \"outlier\" in the public test data. But is this still the case for the private test data?</strong></p>\n<p><strong>My concern is that, under this evaluation metric, if there is one or two similar outliers in the private test data, the public leaderboard is just misleading rather than indicative of the final result…</strong></p>\n<p>Should other metric be considered (one that somehow limits the influence of a single row / compund), so that the public leaderboard and private leaderboard score will be more aligned?</p>",
          "rawMarkdown": "> Perhaps @ryanholbrook can help?\n\nUpon a bit of further investigation. It seems like the in the train data, a compound `MLN 2238` that causes high DEs can \"single-handedly\" contirbutes to a bulk part of the MRRMSE.\n\nSince MRRMSE simply average over rows (instead of somehow limitting the influence of a single row/compound), a single outlier compound like `MLN 2238` can outweight all other compounds combined in its contribution to the MRRMSE in the train data. **Judging from the 0.666 MRRMSE score for the sample submission (all 0), there is no similar such \"outlier\" in the public test data. But is this still the case for the private test data?**\n\n**My concern is that, under this evaluation metric, if there is one or two similar outliers in the private test data, the public leaderboard is just misleading rather than indicative of the final result...**\n\nShould other metric be considered (one that somehow limits the influence of a single row / compund), so that the public leaderboard and private leaderboard score will be more aligned?\n",
          "votes": 4
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2445109,
      "author_name": "Anton Popov",
      "author_url": "",
      "post_date": "2023-09-18T16:14:29.450000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/qihuaz\" target=\"_blank\">@qihuaz</a>!</p>\n<p>The following is the naive implementation of MRRMSE:</p>\n<pre><code> () -&gt; np.float64:\n    \n     Y_true.shape == Y_pred.shape\n    R, n = Y_true.shape\n    result = \n     i  (R):\n        row = \n         j  (n):\n            row += (Y_true[i, j] - Y_pred[i, j]) ** \n        row = np.sqrt(row / n)\n        result += row\n     result / R\n</code></pre>\n<p>But as expected it is slow because of Python's for loop:</p>\n<pre><code>(df_train[target_cols].values, df_train[target_cols].values)\n times: user . s, sys: . ms, total: . s\n time: . s\n</code></pre>\n<p>The vectorized version looks like:</p>\n<pre><code> () -&gt; np.float64:\n    \n     Y_true.shape == Y_pred.shape\n    R, _ = Y_true.shape\n    result = \n     i  (R):\n        row_sum = np.square(Y_true[i, :] - Y_pred[i, :]).mean()\n        result += row_sum\n     result / R\n</code></pre>\n<pre><code>(df_train[target_cols].values, df_train[target_cols].values)\n times: user . ms, sys: . ms, total: . ms\n time: . ms\n</code></pre>\n<p>I hope I understood your issue properly and snippets from above may be of use.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2445344,
          "author_name": "qihuaz",
          "author_url": "",
          "post_date": "2023-09-18T18:50:14.900000",
          "content": "<p>Thanks for the implementation. </p>\n<p>My code was trying to do a quick and dirty calculation of the mrrmse result if 0s are submitted for all prediction. I didn't run yours, but they should come to the same number. </p>\n<p>My issue is that it seems like under this metric, the result is not comparable accross train/public test/private test (please see my reply to <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> for more details).</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2445676,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-09-19T02:57:15.873000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2445678,
          "author_name": "hoho",
          "author_url": "",
          "post_date": "2023-09-19T02:57:55.630000",
          "content": "<p><a href=\"https://www.kaggle.com/calmscout\" target=\"_blank\">@calmscout</a> I think the <code>np.sqrt</code> are missing in the <code>mrrmse_vectorized</code>?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2454047,
          "author_name": "Pablo Rodriguez-Mier",
          "author_url": "",
          "post_date": "2023-09-24T14:25:01.507000",
          "content": "<p>Here is the code more efficent and fixing the missing root:</p>\n<pre><code> () -&gt; np.float64:\n    squared_diffs = np.square(Y_true - Y_pred)\n    row_means = squared_diffs.mean(axis=)\n    root_means = np.sqrt(row_means)\n     root_means.mean()\n</code></pre>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2444970,
      "author_name": "Daniel Burkhardt",
      "author_url": "",
      "post_date": "2023-09-18T15:02:28.303000",
      "content": "<p>Perhaps <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> can help?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2445355,
          "author_name": "qihuaz",
          "author_url": "",
          "post_date": "2023-09-18T18:59:49.860000",
          "content": "<blockquote>\n  <p>Perhaps <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> can help?</p>\n</blockquote>\n<p>Upon a bit of further investigation. It seems like the in the train data, a compound <code>MLN 2238</code> that causes high DEs can \"single-handedly\" contirbutes to a bulk part of the MRRMSE.</p>\n<p>Since MRRMSE simply average over rows (instead of somehow limitting the influence of a single row/compound), a single outlier compound like <code>MLN 2238</code> can outweight all other compounds combined in its contribution to the MRRMSE in the train data. <strong>Judging from the 0.666 MRRMSE score for the sample submission (all 0), there is no similar such \"outlier\" in the public test data. But is this still the case for the private test data?</strong></p>\n<p><strong>My concern is that, under this evaluation metric, if there is one or two similar outliers in the private test data, the public leaderboard is just misleading rather than indicative of the final result…</strong></p>\n<p>Should other metric be considered (one that somehow limits the influence of a single row / compund), so that the public leaderboard and private leaderboard score will be more aligned?</p>",
          "votes": 4,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2443398": "hello all, I am tring to reproduce the evaluation metric Mean Rowwise Root Mean Squared Error (MRRMSE) with the train data. But the result looks much different from the public leaderboard score.\n\nJust as a sanity check, I try to see what the score will be if I predict all 0 for B cells and Myeloid cells, evaluated by the availble train data (excluding positive controls). I am expecting to get something around 0.666 (what you would get by submitting all 0s to the public leaderboard).\n\nSince train is different than test, so I am expecting the value around but not exactly 0.666, maybe in the 0.3~0.9 range. Butinstad I got 2.0345?\n\n**Am I doing something wrong? Or the public test data is just dramastically different from the train data?**\n\nHere's how I got the 2.0345 from the processed train data:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1716193%2F5e6508336e1af45bb6b04fed4404db2d%2FScreenshot%202023-09-17%20133355.png?generation=1694972045873128&alt=media)\nthen\n\n$$\\textrm{MRRMSE} = \\frac{1}{R}\\sum_{i=1}^R\\left(\\frac{1}{n} \\sum_{j=1}^{n} (y_{ij} - \\widehat{y}_{ij})^2\\right)^{1/2}$$\n\n```python\n(\n    (filtered_df_train.iloc[:, 5:] ** 2     # (y_ij - yhat_{ij}) ^ 2\n    ).mean(axis=1) ** 0.5                   # take the mean of n genes then sqaure root\n).mean()                                    # take the mean by R\n```\n\nI hope I am wrong here.. otherwise it would just a lot of guessing game rather than analysis if train set and the test set are distributed so differently... \n\nThe notebook:\nhttps://www.kaggle.com/code/qihuaz/mean-rowwise-root-mean-squared-error?scriptVersionId=143324165",
    "2445109": "Hi @qihuaz!\n\nThe following is the naive implementation of MRRMSE:\n```Python\ndef mrrmse(Y_true:np.ndarray, Y_pred:np.ndarray) -> np.float64:\n    \"\"\"Naive implementation of MRRMSE.\"\"\"\n    assert Y_true.shape == Y_pred.shape\n    R, n = Y_true.shape\n    result = 0\n    for i in range(R):\n        row = 0\n        for j in range(n):\n            row += (Y_true[i, j] - Y_pred[i, j]) ** 2\n        row = np.sqrt(row / n)\n        result += row\n    return result / R\n```\nBut as expected it is slow because of Python's for loop:\n```%%time\nmrrmse(df_train[target_cols].values, df_train[target_cols].values)\nCPU times: user 4.31 s, sys: 4.72 ms, total: 4.32 s\nWall time: 4.32 s\n```\n\nThe vectorized version looks like:\n```Python\ndef mrrmse_vectorized(Y_true:np.ndarray, Y_pred:np.ndarray) -> np.float64:\n    \"\"\"Vectorized implementation of MRRMSE.\"\"\"\n    assert Y_true.shape == Y_pred.shape\n    R, _ = Y_true.shape\n    result = 0\n    for i in range(R):\n        row_sum = np.square(Y_true[i, :] - Y_pred[i, :]).mean()\n        result += row_sum\n    return result / R\n```\n\n```%%time\nmrrmse_vectorized(df_train[target_cols].values, df_train[target_cols].values)\nCPU times: user 67.8 ms, sys: 15.9 ms, total: 83.8 ms\nWall time: 82.6 ms\n```\n\n\nI hope I understood your issue properly and snippets from above may be of use.",
    "2444970": "Perhaps @ryanholbrook can help?"
  }
}