{
  "id": 578290,
  "title": "How do you choose trained models to submit.",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/578290",
  "author_name": "",
  "post_date": "2025-05-09T20:20:28.835647Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>First, after I trained, I don't know which model's performance is better, because the metric I get locally is much differet from the score I get from submission. And second the confidence thresh seems very influence the submission score, and different trained models get best score on different confidence thresh. For examples, with one model, setting confidence thresh to 0.25 can get best score, but with another model setting to 0.65 get the best score.<br>\nSo, how do you make proper submission with only 5 times limits a day? Thanks</p>",
  "messages": [
    {
      "id": "3198713",
      "postDate": "05/09/2025 20:20:28",
      "content": "<p>First, after I trained, I don't know which model's performance is better, because the metric I get locally is much differet from the score I get from submission. And second the confidence thresh seems very influence the submission score, and different trained models get best score on different confidence thresh. For examples, with one model, setting confidence thresh to 0.25 can get best score, but with another model setting to 0.65 get the best score.<br>\nSo, how do you make proper submission with only 5 times limits a day? Thanks</p>",
      "rawMarkdown": "First, after I trained, I don't know which model's performance is better, because the metric I get locally is much differet from the score I get from submission. And second the confidence thresh seems very influence the submission score, and different trained models get best score on different confidence thresh. For examples, with one model, setting confidence thresh to 0.25 can get best score, but with another model setting to 0.65 get the best score.\nSo, how do you make proper submission with only 5 times limits a day? Thanks",
      "votes": null
    },
    {
      "id": "3198754",
      "postDate": "05/09/2025 22:12:38",
      "content": "<p>Use cross validation to get the least noisy experimental results and then you can either ensemble your folds (such as averaging predictions) or just pick a fold that seems to reflect an accurate score.  For the threshold use threshold tuning or even something like non maximum suppression.  </p>\n<p>Do the threshold tuning per model as it will vary based on the training data distribution :).</p>",
      "rawMarkdown": "Use cross validation to get the least noisy experimental results and then you can either ensemble your folds (such as averaging predictions) or just pick a fold that seems to reflect an accurate score.  For the threshold use threshold tuning or even something like non maximum suppression.  \n\nDo the threshold tuning per model as it will vary based on the training data distribution :).",
      "votes": null
    },
    {
      "id": "3202483",
      "postDate": "05/15/2025 13:03:51",
      "content": "<p>Does the cross validation mean that in the submission I should fuse the prediction of every model that get from every fold? And how I judge the model is good or not before I submit to get the real score?</p>",
      "rawMarkdown": "Does the cross validation mean that in the submission I should fuse the prediction of every model that get from every fold? And how I judge the model is good or not before I submit to get the real score?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3198754,
      "author_name": "connorjd",
      "author_url": "",
      "post_date": "05/09/2025 22:12:38",
      "content": "<p>Use cross validation to get the least noisy experimental results and then you can either ensemble your folds (such as averaging predictions) or just pick a fold that seems to reflect an accurate score.  For the threshold use threshold tuning or even something like non maximum suppression.  </p>\n<p>Do the threshold tuning per model as it will vary based on the training data distribution :).</p>",
      "votes": null,
      "replies": [
        {
          "id": 3202483,
          "author_name": "xyzxyzxyzxyzxz",
          "author_url": "",
          "post_date": "05/15/2025 13:03:51",
          "content": "<p>Does the cross validation mean that in the submission I should fuse the prediction of every model that get from every fold? And how I judge the model is good or not before I submit to get the real score?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3198713": "First, after I trained, I don't know which model's performance is better, because the metric I get locally is much differet from the score I get from submission. And second the confidence thresh seems very influence the submission score, and different trained models get best score on different confidence thresh. For examples, with one model, setting confidence thresh to 0.25 can get best score, but with another model setting to 0.65 get the best score.\nSo, how do you make proper submission with only 5 times limits a day? Thanks",
    "3198754": "Use cross validation to get the least noisy experimental results and then you can either ensemble your folds (such as averaging predictions) or just pick a fold that seems to reflect an accurate score.  For the threshold use threshold tuning or even something like non maximum suppression.  \n\nDo the threshold tuning per model as it will vary based on the training data distribution :).",
    "3202483": "Does the cross validation mean that in the submission I should fuse the prediction of every model that get from every fold? And how I judge the model is good or not before I submit to get the real score?"
  },
  "source": "meta"
}