{
  "id": 352655,
  "title": "Negative public score when only evaluating multiome prediction",
  "url": "/competitions/open-problems-multimodal/discussion/352655",
  "author_name": "",
  "post_date": "2022-09-15T08:55:10.908523100Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I wanted to evaluate my prediction for CITE-seq and multiome separately, so I can better understand the performance.</p>\n<p>My model got a score of 0.799 for all cells, and 0.240 for CITE-seq cells, however, for multiome cells, the score is -0.44. See below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F642483%2Fb7460d1a0030f41dd6923c6fc4bbddc3%2FScreenshot%202022-09-15%20at%2010.50.49.png?generation=1663231869661974&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F642483%2F8780d222cbef72f030b63ab790960da3%2FScreenshot%202022-09-15%20at%2010.51.16.png?generation=1663231890215826&amp;alt=media\" alt=\"\"></p>\n<p>What I did is to set the prediction for CITE-seq or multiome cells as zeros, and then combine it with another prediction.</p>\n<pre><code># save all prediction\ndf_eval = pd.concat([df_eval_cite, df_eval_multi], ignore_index=True)\ndf_eval.to_csv(\"./submission.csv\", index=False, float_format='%.3f')\n\n# save only cite-seq predicton by setting multiome as zeros\ndf_eval_multi_zero = df_eval_multi.copy()\ndf_eval_multi_zero['target'] = 0.0\ndf_eval = pd.concat([df_eval_cite, df_eval_multi_zero], ignore_index=True)\ndf_eval.to_csv(\"./submission_cite.csv\", index=False, float_format='%.3f')\n\n# save only multiome predicton by setting multiome as zeros\ndf_eval_cite_zero = df_eval_cite.copy()\ndf_eval_cite_zero['target'] = 0.0\ndf_eval = pd.concat([df_eval_cite_zero, df_eval_multi], ignore_index=True)\ndf_eval.to_csv(\"./submission_multi.csv\", index=False, float_format='%.3f')\n</code></pre>\n<p>Does any have such an issue?</p>",
  "messages": [
    {
      "id": "1940259",
      "postDate": "09/15/2022 08:55:10",
      "content": "<p>I wanted to evaluate my prediction for CITE-seq and multiome separately, so I can better understand the performance.</p>\n<p>My model got a score of 0.799 for all cells, and 0.240 for CITE-seq cells, however, for multiome cells, the score is -0.44. See below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F642483%2Fb7460d1a0030f41dd6923c6fc4bbddc3%2FScreenshot%202022-09-15%20at%2010.50.49.png?generation=1663231869661974&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F642483%2F8780d222cbef72f030b63ab790960da3%2FScreenshot%202022-09-15%20at%2010.51.16.png?generation=1663231890215826&amp;alt=media\" alt=\"\"></p>\n<p>What I did is to set the prediction for CITE-seq or multiome cells as zeros, and then combine it with another prediction.</p>\n<pre><code># save all prediction\ndf_eval = pd.concat([df_eval_cite, df_eval_multi], ignore_index=True)\ndf_eval.to_csv(\"./submission.csv\", index=False, float_format='%.3f')\n\n# save only cite-seq predicton by setting multiome as zeros\ndf_eval_multi_zero = df_eval_multi.copy()\ndf_eval_multi_zero['target'] = 0.0\ndf_eval = pd.concat([df_eval_cite, df_eval_multi_zero], ignore_index=True)\ndf_eval.to_csv(\"./submission_cite.csv\", index=False, float_format='%.3f')\n\n# save only multiome predicton by setting multiome as zeros\ndf_eval_cite_zero = df_eval_cite.copy()\ndf_eval_cite_zero['target'] = 0.0\ndf_eval = pd.concat([df_eval_cite_zero, df_eval_multi], ignore_index=True)\ndf_eval.to_csv(\"./submission_multi.csv\", index=False, float_format='%.3f')\n</code></pre>\n<p>Does any have such an issue?</p>",
      "rawMarkdown": "I wanted to evaluate my prediction for CITE-seq and multiome separately, so I can better understand the performance.\n\nMy model got a score of 0.799 for all cells, and 0.240 for CITE-seq cells, however, for multiome cells, the score is -0.44. See below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F642483%2Fb7460d1a0030f41dd6923c6fc4bbddc3%2FScreenshot%202022-09-15%20at%2010.50.49.png?generation=1663231869661974&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F642483%2F8780d222cbef72f030b63ab790960da3%2FScreenshot%202022-09-15%20at%2010.51.16.png?generation=1663231890215826&alt=media)\n\nWhat I did is to set the prediction for CITE-seq or multiome cells as zeros, and then combine it with another prediction.\n\n```\n# save all prediction\ndf_eval = pd.concat([df_eval_cite, df_eval_multi], ignore_index=True)\ndf_eval.to_csv(\"./submission.csv\", index=False, float_format='%.3f')\n\n# save only cite-seq predicton by setting multiome as zeros\ndf_eval_multi_zero = df_eval_multi.copy()\ndf_eval_multi_zero['target'] = 0.0\ndf_eval = pd.concat([df_eval_cite, df_eval_multi_zero], ignore_index=True)\ndf_eval.to_csv(\"./submission_cite.csv\", index=False, float_format='%.3f')\n\n# save only multiome predicton by setting multiome as zeros\ndf_eval_cite_zero = df_eval_cite.copy()\ndf_eval_cite_zero['target'] = 0.0\ndf_eval = pd.concat([df_eval_cite_zero, df_eval_multi], ignore_index=True)\ndf_eval.to_csv(\"./submission_multi.csv\", index=False, float_format='%.3f')\n```\n\nDoes any have such an issue?",
      "votes": null
    },
    {
      "id": "1940327",
      "postDate": "09/15/2022 09:39:05",
      "content": "<p>from evaluation page:</p>\n<p>\"If a sample's predictions are all the same, the correlation for that sample is scored as -1.0\"</p>",
      "rawMarkdown": "from evaluation page:\n\n\"If a sample's predictions are all the same, the correlation for that sample is scored as -1.0\"",
      "votes": null
    },
    {
      "id": "1940331",
      "postDate": "09/15/2022 09:43:15",
      "content": "<p>Ohh, didn't notice that. <br>\nThanks</p>",
      "rawMarkdown": "Ohh, didn't notice that. \nThanks",
      "votes": null
    },
    {
      "id": "1940902",
      "postDate": "09/15/2022 16:10:56",
      "content": "<p>The leaderboard is calculated with approximately 42% of the test data. The final results will be based on the other 58%. So, you could evaluate both the multiome part and Cite part only on the Public leaderboard,  however, choose the random submission for the first step, and change each part to evaluate both Multiome and Cite part, with taking attention to your result will represent 42% of the test data.</p>",
      "rawMarkdown": "The leaderboard is calculated with approximately 42% of the test data. The final results will be based on the other 58%. So, you could evaluate both the multiome part and Cite part only on the Public leaderboard,  however, choose the random submission for the first step, and change each part to evaluate both Multiome and Cite part, with taking attention to your result will represent 42% of the test data.",
      "votes": null
    },
    {
      "id": "1953786",
      "postDate": "09/24/2022 18:20:42",
      "content": "<p>I got the -.44 score also when I looked at just the Multiome data with the CITE data that was merged set at all zeros (unintentionally!) </p>",
      "rawMarkdown": "I got the -.44 score also when I looked at just the Multiome data with the CITE data that was merged set at all zeros (unintentionally!)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1940327,
      "author_name": "julesvanligtenberg",
      "author_url": "",
      "post_date": "09/15/2022 09:39:05",
      "content": "<p>from evaluation page:</p>\n<p>\"If a sample's predictions are all the same, the correlation for that sample is scored as -1.0\"</p>",
      "votes": null,
      "replies": [
        {
          "id": 1940331,
          "author_name": "zhijianli",
          "author_url": "",
          "post_date": "09/15/2022 09:43:15",
          "content": "<p>Ohh, didn't notice that. <br>\nThanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1940902,
      "author_name": "youneseloiarm",
      "author_url": "",
      "post_date": "09/15/2022 16:10:56",
      "content": "<p>The leaderboard is calculated with approximately 42% of the test data. The final results will be based on the other 58%. So, you could evaluate both the multiome part and Cite part only on the Public leaderboard,  however, choose the random submission for the first step, and change each part to evaluate both Multiome and Cite part, with taking attention to your result will represent 42% of the test data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1953786,
      "author_name": "bkosar1640",
      "author_url": "",
      "post_date": "09/24/2022 18:20:42",
      "content": "<p>I got the -.44 score also when I looked at just the Multiome data with the CITE data that was merged set at all zeros (unintentionally!) </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1940259": "I wanted to evaluate my prediction for CITE-seq and multiome separately, so I can better understand the performance.\n\nMy model got a score of 0.799 for all cells, and 0.240 for CITE-seq cells, however, for multiome cells, the score is -0.44. See below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F642483%2Fb7460d1a0030f41dd6923c6fc4bbddc3%2FScreenshot%202022-09-15%20at%2010.50.49.png?generation=1663231869661974&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F642483%2F8780d222cbef72f030b63ab790960da3%2FScreenshot%202022-09-15%20at%2010.51.16.png?generation=1663231890215826&alt=media)\n\nWhat I did is to set the prediction for CITE-seq or multiome cells as zeros, and then combine it with another prediction.\n\n```\n# save all prediction\ndf_eval = pd.concat([df_eval_cite, df_eval_multi], ignore_index=True)\ndf_eval.to_csv(\"./submission.csv\", index=False, float_format='%.3f')\n\n# save only cite-seq predicton by setting multiome as zeros\ndf_eval_multi_zero = df_eval_multi.copy()\ndf_eval_multi_zero['target'] = 0.0\ndf_eval = pd.concat([df_eval_cite, df_eval_multi_zero], ignore_index=True)\ndf_eval.to_csv(\"./submission_cite.csv\", index=False, float_format='%.3f')\n\n# save only multiome predicton by setting multiome as zeros\ndf_eval_cite_zero = df_eval_cite.copy()\ndf_eval_cite_zero['target'] = 0.0\ndf_eval = pd.concat([df_eval_cite_zero, df_eval_multi], ignore_index=True)\ndf_eval.to_csv(\"./submission_multi.csv\", index=False, float_format='%.3f')\n```\n\nDoes any have such an issue?",
    "1940327": "from evaluation page:\n\n\"If a sample's predictions are all the same, the correlation for that sample is scored as -1.0\"",
    "1940331": "Ohh, didn't notice that. \nThanks",
    "1940902": "The leaderboard is calculated with approximately 42% of the test data. The final results will be based on the other 58%. So, you could evaluate both the multiome part and Cite part only on the Public leaderboard,  however, choose the random submission for the first step, and change each part to evaluate both Multiome and Cite part, with taking attention to your result will represent 42% of the test data.",
    "1953786": "I got the -.44 score also when I looked at just the Multiome data with the CITE data that was merged set at all zeros (unintentionally!)"
  },
  "source": "meta"
}