{
  "id": 515684,
  "title": "Evaluation and Submission File",
  "url": "/competitions/leash-BELKA/discussion/515684",
  "author_name": "",
  "post_date": "2024-06-29T11:30:13.097488100Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Sure, here is the paragraph rewritten in a clearer and more professional manner:</p>\n<hr>\n<h3>Discussion and Request for Assistance</h3>\n<p><strong>Evaluation:</strong>  <br>\nThe metric for this competition is the average precision calculated for each (protein, split group) and then averaged for the final score. Please refer to this <a href=\"#\">forum post</a> for important details.</p>\n<p><strong>Submission File:</strong>  <br>\nFor each id in the test set, you must predict a probability for the binary target \"binds the target\". The submission file should contain a header and follow this format:</p>\n<pre><code>id,binds\n.\n.\n.\n...\n</code></pre>\n<p>Our team has developed a model that outputs a predicted value of 0 or 1, indicating \"no bind\" or \"bind\" respectively. However, we are confused by the submission file format, which specifies a value of 0.5. We understand this represents the average predicted value for each protein and split group. We believe this averaging method might yield more accurate results for the unknown SMILES (new library). Could anyone provide further clarification on how to appropriately format our submission file to align with the competition requirements? Any insights or guidance would be greatly appreciated.<br>\nthe questions here are if we submit the file without taking the average will be suitable and if not how can we make this split group and 0.5 as value? <br>\nlast one: is 0, 0.5, or 1 the only values allowed or we can get any number between (0- 1)</p>",
  "messages": [
    {
      "id": "2895818",
      "postDate": "06/29/2024 11:30:13",
      "content": "<p>Sure, here is the paragraph rewritten in a clearer and more professional manner:</p>\n<hr>\n<h3>Discussion and Request for Assistance</h3>\n<p><strong>Evaluation:</strong>  <br>\nThe metric for this competition is the average precision calculated for each (protein, split group) and then averaged for the final score. Please refer to this <a href=\"#\">forum post</a> for important details.</p>\n<p><strong>Submission File:</strong>  <br>\nFor each id in the test set, you must predict a probability for the binary target \"binds the target\". The submission file should contain a header and follow this format:</p>\n<pre><code>id,binds\n.\n.\n.\n...\n</code></pre>\n<p>Our team has developed a model that outputs a predicted value of 0 or 1, indicating \"no bind\" or \"bind\" respectively. However, we are confused by the submission file format, which specifies a value of 0.5. We understand this represents the average predicted value for each protein and split group. We believe this averaging method might yield more accurate results for the unknown SMILES (new library). Could anyone provide further clarification on how to appropriately format our submission file to align with the competition requirements? Any insights or guidance would be greatly appreciated.<br>\nthe questions here are if we submit the file without taking the average will be suitable and if not how can we make this split group and 0.5 as value? <br>\nlast one: is 0, 0.5, or 1 the only values allowed or we can get any number between (0- 1)</p>",
      "rawMarkdown": "Sure, here is the paragraph rewritten in a clearer and more professional manner:\n\n---\n\n### Discussion and Request for Assistance\n\n**Evaluation:**  \nThe metric for this competition is the average precision calculated for each (protein, split group) and then averaged for the final score. Please refer to this [forum post](#) for important details.\n\n**Submission File:**  \nFor each id in the test set, you must predict a probability for the binary target \"binds the target\". The submission file should contain a header and follow this format:\n\n```\nid,binds\n295246830,0.5\n295246831,0.5\n295246832,0.5\n...\n```\n\nOur team has developed a model that outputs a predicted value of 0 or 1, indicating \"no bind\" or \"bind\" respectively. However, we are confused by the submission file format, which specifies a value of 0.5. We understand this represents the average predicted value for each protein and split group. We believe this averaging method might yield more accurate results for the unknown SMILES (new library). Could anyone provide further clarification on how to appropriately format our submission file to align with the competition requirements? Any insights or guidance would be greatly appreciated.\nthe questions here are if we submit the file without taking the average will be suitable and if not how can we make this split group and 0.5 as value? \nlast one: is 0, 0.5, or 1 the only values allowed or we can get any number between (0- 1)",
      "votes": null
    },
    {
      "id": "2902927",
      "postDate": "07/03/2024 14:37:48",
      "content": "<p>The 0.5 in the sample submission is just a placeholder, it means <em>Your prediction goes here</em>. The submission file should be identical in format to the sample submission file, except that the predictions are hopefully not all 0.5. You can predict any number between 0.0 and 1.0 inclusive.</p>\n<p>Probabilistic predictions are recommended. Models that only make predictions of 0.0 or 1.0 are likely to score poorly, since they don't discriminate well between the likelihoods of different compounds being binders. I think it's helpful to think of your probabilities as a ranking of compounds from the least likely to the most likely to bind. </p>",
      "rawMarkdown": "The 0.5 in the sample submission is just a placeholder, it means *Your prediction goes here*. The submission file should be identical in format to the sample submission file, except that the predictions are hopefully not all 0.5. You can predict any number between 0.0 and 1.0 inclusive.\n\nProbabilistic predictions are recommended. Models that only make predictions of 0.0 or 1.0 are likely to score poorly, since they don't discriminate well between the likelihoods of different compounds being binders. I think it's helpful to think of your probabilities as a ranking of compounds from the least likely to the most likely to bind.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2902927,
      "author_name": "jbomitchell",
      "author_url": "",
      "post_date": "07/03/2024 14:37:48",
      "content": "<p>The 0.5 in the sample submission is just a placeholder, it means <em>Your prediction goes here</em>. The submission file should be identical in format to the sample submission file, except that the predictions are hopefully not all 0.5. You can predict any number between 0.0 and 1.0 inclusive.</p>\n<p>Probabilistic predictions are recommended. Models that only make predictions of 0.0 or 1.0 are likely to score poorly, since they don't discriminate well between the likelihoods of different compounds being binders. I think it's helpful to think of your probabilities as a ranking of compounds from the least likely to the most likely to bind. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2895818": "Sure, here is the paragraph rewritten in a clearer and more professional manner:\n\n---\n\n### Discussion and Request for Assistance\n\n**Evaluation:**  \nThe metric for this competition is the average precision calculated for each (protein, split group) and then averaged for the final score. Please refer to this [forum post](#) for important details.\n\n**Submission File:**  \nFor each id in the test set, you must predict a probability for the binary target \"binds the target\". The submission file should contain a header and follow this format:\n\n```\nid,binds\n295246830,0.5\n295246831,0.5\n295246832,0.5\n...\n```\n\nOur team has developed a model that outputs a predicted value of 0 or 1, indicating \"no bind\" or \"bind\" respectively. However, we are confused by the submission file format, which specifies a value of 0.5. We understand this represents the average predicted value for each protein and split group. We believe this averaging method might yield more accurate results for the unknown SMILES (new library). Could anyone provide further clarification on how to appropriately format our submission file to align with the competition requirements? Any insights or guidance would be greatly appreciated.\nthe questions here are if we submit the file without taking the average will be suitable and if not how can we make this split group and 0.5 as value? \nlast one: is 0, 0.5, or 1 the only values allowed or we can get any number between (0- 1)",
    "2902927": "The 0.5 in the sample submission is just a placeholder, it means *Your prediction goes here*. The submission file should be identical in format to the sample submission file, except that the predictions are hopefully not all 0.5. You can predict any number between 0.0 and 1.0 inclusive.\n\nProbabilistic predictions are recommended. Models that only make predictions of 0.0 or 1.0 are likely to score poorly, since they don't discriminate well between the likelihoods of different compounds being binders. I think it's helpful to think of your probabilities as a ranking of compounds from the least likely to the most likely to bind."
  },
  "source": "meta"
}