{
  "id": 581111,
  "title": "A weight solver for ensembling to maximize corr",
  "url": "/competitions/drw-crypto-market-prediction/discussion/581111",
  "author_name": "Yourui Wang",
  "post_date": "2025-05-28T11:33:46.525000",
  "votes": 19,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>I'm here to share a function to give best weights for your multiple models to ensemble to generate maximized correlation coefficient.</p>\n<pre><code> ():\n    n = cov_mat.shape[] - \n    Sigma_XX = cov_mat[:n, :n]\n    Sigma_Xy = cov_mat[:n, n].reshape(-, )\n    Var_y = cov_mat[n, n]\n    w = np.linalg.solve(Sigma_XX, Sigma_Xy)\n    norm = np.sqrt(w.T @ Sigma_XX @ w)[, ]\n    w_normalized = w / norm\n     w_normalized.flatten()\n</code></pre>\n<p>The input <code>cov_mat</code> should be your covariance matrix between your multiple model outputs and label (Note: the last one row and column represent label).</p>\n<p>The output of the function is the optimal weights for your model outputs. You can simply use it to generate your final prediction.</p>\n<p>This can only be used in train_data. However, it should give a good estimate about proper weights between your models. You can compare with ridge weights and other methods.</p>",
  "messages": [
    {
      "id": 3211397,
      "postDate": "2025-05-28T11:33:46.527Z",
      "content": "<p>Hi everyone,</p>\n<p>I'm here to share a function to give best weights for your multiple models to ensemble to generate maximized correlation coefficient.</p>\n<pre><code> ():\n    n = cov_mat.shape[] - \n    Sigma_XX = cov_mat[:n, :n]\n    Sigma_Xy = cov_mat[:n, n].reshape(-, )\n    Var_y = cov_mat[n, n]\n    w = np.linalg.solve(Sigma_XX, Sigma_Xy)\n    norm = np.sqrt(w.T @ Sigma_XX @ w)[, ]\n    w_normalized = w / norm\n     w_normalized.flatten()\n</code></pre>\n<p>The input <code>cov_mat</code> should be your covariance matrix between your multiple model outputs and label (Note: the last one row and column represent label).</p>\n<p>The output of the function is the optimal weights for your model outputs. You can simply use it to generate your final prediction.</p>\n<p>This can only be used in train_data. However, it should give a good estimate about proper weights between your models. You can compare with ridge weights and other methods.</p>",
      "rawMarkdown": "Hi everyone,\n\nI'm here to share a function to give best weights for your multiple models to ensemble to generate maximized correlation coefficient.\n\n```python\ndef max_corr_projection(cov_mat):\n    n = cov_mat.shape[0] - 1\n    Sigma_XX = cov_mat[:n, :n]\n    Sigma_Xy = cov_mat[:n, n].reshape(-1, 1)\n    Var_y = cov_mat[n, n]\n    w = np.linalg.solve(Sigma_XX, Sigma_Xy)\n    norm = np.sqrt(w.T @ Sigma_XX @ w)[0, 0]\n    w_normalized = w / norm\n    return w_normalized.flatten()\n```\n\nThe input `cov_mat` should be your covariance matrix between your multiple model outputs and label (Note: the last one row and column represent label).\n\nThe output of the function is the optimal weights for your model outputs. You can simply use it to generate your final prediction.\n\nThis can only be used in train_data. However, it should give a good estimate about proper weights between your models. You can compare with ridge weights and other methods.\n",
      "votes": 19
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3211397": "Hi everyone,\n\nI'm here to share a function to give best weights for your multiple models to ensemble to generate maximized correlation coefficient.\n\n```python\ndef max_corr_projection(cov_mat):\n    n = cov_mat.shape[0] - 1\n    Sigma_XX = cov_mat[:n, :n]\n    Sigma_Xy = cov_mat[:n, n].reshape(-1, 1)\n    Var_y = cov_mat[n, n]\n    w = np.linalg.solve(Sigma_XX, Sigma_Xy)\n    norm = np.sqrt(w.T @ Sigma_XX @ w)[0, 0]\n    w_normalized = w / norm\n    return w_normalized.flatten()\n```\n\nThe input `cov_mat` should be your covariance matrix between your multiple model outputs and label (Note: the last one row and column represent label).\n\nThe output of the function is the optimal weights for your model outputs. You can simply use it to generate your final prediction.\n\nThis can only be used in train_data. However, it should give a good estimate about proper weights between your models. You can compare with ridge weights and other methods.\n"
  }
}