{
  "id": 550708,
  "title": "If Your CV Score is Way Off Your LB Score...",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/550708",
  "author_name": "",
  "post_date": "2024-12-09T06:04:02.133797400Z",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>And if you're relying on a function called compute_lb() to calculate your CV score.  Comment out the line that sets valid_id like so:</p>\n<pre><code>def (submit_df, overlay_dir):\n    #valid_id = (submit_df[].())\n    (valid_id)\n</code></pre>\n<p>This works because valid_id is set globally above to:</p>\n<pre><code> = [, ] \n</code></pre>\n<p>If it's not, then add that line immediately below the one you commented out.</p>",
  "messages": [
    {
      "id": "3067281",
      "postDate": "12/09/2024 06:04:02",
      "content": "<p>And if you're relying on a function called compute_lb() to calculate your CV score.  Comment out the line that sets valid_id like so:</p>\n<pre><code>def (submit_df, overlay_dir):\n    #valid_id = (submit_df[].())\n    (valid_id)\n</code></pre>\n<p>This works because valid_id is set globally above to:</p>\n<pre><code> = [, ] \n</code></pre>\n<p>If it's not, then add that line immediately below the one you commented out.</p>",
      "rawMarkdown": "And if you're relying on a function called compute_lb() to calculate your CV score.  Comment out the line that sets valid_id like so:\n\n```\ndef compute_lb(submit_df, overlay_dir):\n    #valid_id = list(submit_df['experiment'].unique())\n    print(valid_id)\n```\n\nThis works because valid_id is set globally above to:\n\n```\nvalid_id = ['TS_6_4', ] # Or whatever you're using for validation.\n```\n\nIf it's not, then add that line immediately below the one you commented out.",
      "votes": null
    },
    {
      "id": "3067461",
      "postDate": "12/09/2024 10:13:16",
      "content": "<p>Hello, I am not sure I follow, isn't the commented line getting the unique experiments from your submission df, essentially 'adapting' to your validation experiment(s)? It works with multiple experiments at once and uses the experiment name to extract the ground truth as well for that experiment, wouldn't it return close to 0 values if it read ground truth values from a different experiment?</p>",
      "rawMarkdown": "Hello, I am not sure I follow, isn't the commented line getting the unique experiments from your submission df, essentially 'adapting' to your validation experiment(s)? It works with multiple experiments at once and uses the experiment name to extract the ground truth as well for that experiment, wouldn't it return close to 0 values if it read ground truth values from a different experiment?",
      "votes": null
    },
    {
      "id": "3067493",
      "postDate": "12/09/2024 10:46:25",
      "content": "<p>The notebook that's taken from (the most popular for the contest so far) generates answers for TS_6_4, TS_5_4, and TS_69_2 even though it trains on both TS_5_4 and TS_69_2.  Only TS_6_4 is used for validation.  Because it's only using TS_6_4 for validation, it should only be calculating the CV score from TS_6_4.  Because it also includes TS_5_4 and TS_69_2 the CV score tends to be elevated (typically about 0.1 too high.)</p>",
      "rawMarkdown": "The notebook that's taken from (the most popular for the contest so far) generates answers for TS_6_4, TS_5_4, and TS_69_2 even though it trains on both TS_5_4 and TS_69_2.  Only TS_6_4 is used for validation.  Because it's only using TS_6_4 for validation, it should only be calculating the CV score from TS_6_4.  Because it also includes TS_5_4 and TS_69_2 the CV score tends to be elevated (typically about 0.1 too high.)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3067461,
      "author_name": "andreizamfir",
      "author_url": "",
      "post_date": "12/09/2024 10:13:16",
      "content": "<p>Hello, I am not sure I follow, isn't the commented line getting the unique experiments from your submission df, essentially 'adapting' to your validation experiment(s)? It works with multiple experiments at once and uses the experiment name to extract the ground truth as well for that experiment, wouldn't it return close to 0 values if it read ground truth values from a different experiment?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3067493,
          "author_name": "davidlist",
          "author_url": "",
          "post_date": "12/09/2024 10:46:25",
          "content": "<p>The notebook that's taken from (the most popular for the contest so far) generates answers for TS_6_4, TS_5_4, and TS_69_2 even though it trains on both TS_5_4 and TS_69_2.  Only TS_6_4 is used for validation.  Because it's only using TS_6_4 for validation, it should only be calculating the CV score from TS_6_4.  Because it also includes TS_5_4 and TS_69_2 the CV score tends to be elevated (typically about 0.1 too high.)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3067281": "And if you're relying on a function called compute_lb() to calculate your CV score.  Comment out the line that sets valid_id like so:\n\n```\ndef compute_lb(submit_df, overlay_dir):\n    #valid_id = list(submit_df['experiment'].unique())\n    print(valid_id)\n```\n\nThis works because valid_id is set globally above to:\n\n```\nvalid_id = ['TS_6_4', ] # Or whatever you're using for validation.\n```\n\nIf it's not, then add that line immediately below the one you commented out.",
    "3067461": "Hello, I am not sure I follow, isn't the commented line getting the unique experiments from your submission df, essentially 'adapting' to your validation experiment(s)? It works with multiple experiments at once and uses the experiment name to extract the ground truth as well for that experiment, wouldn't it return close to 0 values if it read ground truth values from a different experiment?",
    "3067493": "The notebook that's taken from (the most popular for the contest so far) generates answers for TS_6_4, TS_5_4, and TS_69_2 even though it trains on both TS_5_4 and TS_69_2.  Only TS_6_4 is used for validation.  Because it's only using TS_6_4 for validation, it should only be calculating the CV score from TS_6_4.  Because it also includes TS_5_4 and TS_69_2 the CV score tends to be elevated (typically about 0.1 too high.)"
  },
  "source": "meta"
}