{
  "id": 543838,
  "title": "Any Luck Finding a CV Strategy That Aligns with the LB?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/543838",
  "author_name": "Ayman Allawi",
  "post_date": "2024-11-01T18:05:29.932000",
  "votes": 8,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I've tried several CV strategies but haven't found one that aligns well with the LB. For example, I'll make a change that improves performance on CV, but when I test it on the LB, there's no improvement. I'm curious if anyone has found a CV approach that aligns closely with the LB, or if it's just very challenging to find one due to the nature of the data. Any insights would be appreciated!</p>",
  "messages": [
    {
      "id": 3034156,
      "postDate": "2024-11-01T20:36:30.047Z",
      "content": "<p>Yes, it's possible. But not like normal competitions you just need one single validation set to validate your model on and hopefully align your cv score with LB, in this competition you may need test your model's robustness in multiple validation sets and try to find some parttern out of it.</p>",
      "rawMarkdown": "Yes, it's possible. But not like normal competitions you just need one single validation set to validate your model on and hopefully align your cv score with LB, in this competition you may need test your model's robustness in multiple validation sets and try to find some parttern out of it.",
      "votes": 7,
      "replies": [
        {
          "id": 3034185,
          "postDate": "2024-11-01T21:26:41.963Z",
          "content": "<p>I'm not going to ask about your CV strategy, as a successful CV approach is, of course, an advantage that no one would easily share. I just want to know if your CV aligns perfectly with the LB — meaning that everything that improves CV also improves LB. Or have you sometimes had to choose to trust one over the other?</p>",
          "rawMarkdown": "I'm not going to ask about your CV strategy, as a successful CV approach is, of course, an advantage that no one would easily share. I just want to know if your CV aligns perfectly with the LB — meaning that everything that improves CV also improves LB. Or have you sometimes had to choose to trust one over the other?",
          "votes": 2,
          "replies": [
            {
              "id": 3034191,
              "postDate": "2024-11-01T21:31:27.433Z",
              "content": "<p>For now, yes, I can see the local CV and LB align well. </p>",
              "rawMarkdown": "For now, yes, I can see the local CV and LB align well. ",
              "votes": 3
            },
            {
              "id": 3034238,
              "postDate": "2024-11-01T22:54:27.747Z",
              "content": "<p>Thank you for your hints. But in this way are you not risking to overfit to the LB? How do you know that the private will behave like the LB? </p>",
              "rawMarkdown": "Thank you for your hints. But in this way are you not risking to overfit to the LB? How do you know that the private will behave like the LB? ",
              "votes": 1
            },
            {
              "id": 3034244,
              "postDate": "2024-11-01T23:07:14.403Z",
              "content": "<p>I only submit to PB when I see the action I took steadily improve my local CV (except of some submissions for testing purpose). And the CV score from what I see is aligned with the PB I get. </p>",
              "rawMarkdown": "I only submit to PB when I see the action I took steadily improve my local CV (except of some submissions for testing purpose). And the CV score from what I see is aligned with the PB I get. ",
              "votes": 1
            },
            {
              "id": 3034410,
              "postDate": "2024-11-02T05:49:35.350Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 3034048,
      "postDate": "2024-11-01T18:05:29.933Z",
      "content": "<p>I've tried several CV strategies but haven't found one that aligns well with the LB. For example, I'll make a change that improves performance on CV, but when I test it on the LB, there's no improvement. I'm curious if anyone has found a CV approach that aligns closely with the LB, or if it's just very challenging to find one due to the nature of the data. Any insights would be appreciated!</p>",
      "rawMarkdown": "I've tried several CV strategies but haven't found one that aligns well with the LB. For example, I'll make a change that improves performance on CV, but when I test it on the LB, there's no improvement. I'm curious if anyone has found a CV approach that aligns closely with the LB, or if it's just very challenging to find one due to the nature of the data. Any insights would be appreciated!\n\n",
      "votes": 8
    },
    {
      "id": 3034078,
      "postDate": "2024-11-01T18:44:38.213Z",
      "content": "<p>Same here. Tried multiple cv strategies (GroupTimeSeriesSplit, PurgedGroupTimeSeriesSplit, GroupKFold, PurgedGroupKFold, CombinatorialPurgedGroupKFold) to build distributions of learning curves (num_iterations vs competition metric). Then i submitted same model but with different num_iterations (10, 20, 50, 100, 300) to approximate a learning curve for LB. The approximated LB learning curve is always out of the different cv distributions.</p>",
      "rawMarkdown": "Same here. Tried multiple cv strategies (GroupTimeSeriesSplit, PurgedGroupTimeSeriesSplit, GroupKFold, PurgedGroupKFold, CombinatorialPurgedGroupKFold) to build distributions of learning curves (num_iterations vs competition metric). Then i submitted same model but with different num_iterations (10, 20, 50, 100, 300) to approximate a learning curve for LB. The approximated LB learning curve is always out of the different cv distributions.",
      "votes": 1,
      "replies": [
        {
          "id": 3034084,
          "postDate": "2024-11-01T18:53:33.973Z",
          "content": "<p>According to what I saw so far, I believe working based on LB will lead to overfitting and trusting the LB score might lead to significant drop in rank on the private LB </p>",
          "rawMarkdown": "According to what I saw so far, I believe working based on LB will lead to overfitting and trusting the LB score might lead to significant drop in rank on the private LB "
        },
        {
          "id": 3034456,
          "postDate": "2024-11-02T06:51:15.907Z",
          "content": "<p>BTW, are you using symbol_id and time_id as training features?</p>",
          "rawMarkdown": "BTW, are you using symbol_id and time_id as training features?",
          "replies": [
            {
              "id": 3036395,
              "postDate": "2024-11-04T14:25:15.010Z",
              "content": "<p>I'm using time_id but no symbol_id as feature.</p>",
              "rawMarkdown": "I'm using time_id but no symbol_id as feature."
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3034156,
      "author_name": "HAO",
      "author_url": "",
      "post_date": "2024-11-01T20:36:30.047000",
      "content": "<p>Yes, it's possible. But not like normal competitions you just need one single validation set to validate your model on and hopefully align your cv score with LB, in this competition you may need test your model's robustness in multiple validation sets and try to find some parttern out of it.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 3034185,
          "author_name": "Ayman Allawi",
          "author_url": "",
          "post_date": "2024-11-01T21:26:41.963000",
          "content": "<p>I'm not going to ask about your CV strategy, as a successful CV approach is, of course, an advantage that no one would easily share. I just want to know if your CV aligns perfectly with the LB — meaning that everything that improves CV also improves LB. Or have you sometimes had to choose to trust one over the other?</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3034191,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2024-11-01T21:31:27.433000",
              "content": "<p>For now, yes, I can see the local CV and LB align well. </p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3034238,
              "author_name": "gatsu",
              "author_url": "",
              "post_date": "2024-11-01T22:54:27.747000",
              "content": "<p>Thank you for your hints. But in this way are you not risking to overfit to the LB? How do you know that the private will behave like the LB? </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3034244,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2024-11-01T23:07:14.403000",
              "content": "<p>I only submit to PB when I see the action I took steadily improve my local CV (except of some submissions for testing purpose). And the CV score from what I see is aligned with the PB I get. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3034410,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-11-02T05:49:35.350000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3034078,
      "author_name": "gatsu",
      "author_url": "",
      "post_date": "2024-11-01T18:44:38.213000",
      "content": "<p>Same here. Tried multiple cv strategies (GroupTimeSeriesSplit, PurgedGroupTimeSeriesSplit, GroupKFold, PurgedGroupKFold, CombinatorialPurgedGroupKFold) to build distributions of learning curves (num_iterations vs competition metric). Then i submitted same model but with different num_iterations (10, 20, 50, 100, 300) to approximate a learning curve for LB. The approximated LB learning curve is always out of the different cv distributions.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3034084,
          "author_name": "Ayman Allawi",
          "author_url": "",
          "post_date": "2024-11-01T18:53:33.973000",
          "content": "<p>According to what I saw so far, I believe working based on LB will lead to overfitting and trusting the LB score might lead to significant drop in rank on the private LB </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3034456,
          "author_name": "Ayman Allawi",
          "author_url": "",
          "post_date": "2024-11-02T06:51:15.907000",
          "content": "<p>BTW, are you using symbol_id and time_id as training features?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3036395,
              "author_name": "gatsu",
              "author_url": "",
              "post_date": "2024-11-04T14:25:15.010000",
              "content": "<p>I'm using time_id but no symbol_id as feature.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3034156": "Yes, it's possible. But not like normal competitions you just need one single validation set to validate your model on and hopefully align your cv score with LB, in this competition you may need test your model's robustness in multiple validation sets and try to find some parttern out of it.",
    "3034048": "I've tried several CV strategies but haven't found one that aligns well with the LB. For example, I'll make a change that improves performance on CV, but when I test it on the LB, there's no improvement. I'm curious if anyone has found a CV approach that aligns closely with the LB, or if it's just very challenging to find one due to the nature of the data. Any insights would be appreciated!\n\n",
    "3034078": "Same here. Tried multiple cv strategies (GroupTimeSeriesSplit, PurgedGroupTimeSeriesSplit, GroupKFold, PurgedGroupKFold, CombinatorialPurgedGroupKFold) to build distributions of learning curves (num_iterations vs competition metric). Then i submitted same model but with different num_iterations (10, 20, 50, 100, 300) to approximate a learning curve for LB. The approximated LB learning curve is always out of the different cv distributions."
  }
}