{
  "id": 173587,
  "title": "Cross Validation with Randomly Cropped Segments for Every Fold?",
  "url": "/competitions/birdsong-recognition/discussion/173587",
  "author_name": "",
  "post_date": "2020-08-09T22:55:17.893917Z",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I've noticed most of the public kernels are using randomly cropped to 5 seconds segments, use them as a training for every epoch and do this with 5 fold. I am wondering would the CV value is credible in that case? I think the data samples for each fold has different trining data for the same label and it may cause the CV value unstable. Maybe caching the cropped segments for each epoch and run the model would give the stable CVs. <br>\nOn the other side in my mind, it is okay with current one because it is random! The expectation of amount of information we get from the training data in each fold would be similar.<br>\nHow do you think?</p>\n<p>Thank you</p>",
  "messages": [
    {
      "id": "964464",
      "postDate": "08/09/2020 22:55:17",
      "content": "<p>I've noticed most of the public kernels are using randomly cropped to 5 seconds segments, use them as a training for every epoch and do this with 5 fold. I am wondering would the CV value is credible in that case? I think the data samples for each fold has different trining data for the same label and it may cause the CV value unstable. Maybe caching the cropped segments for each epoch and run the model would give the stable CVs. <br>\nOn the other side in my mind, it is okay with current one because it is random! The expectation of amount of information we get from the training data in each fold would be similar.<br>\nHow do you think?</p>\n<p>Thank you</p>",
      "rawMarkdown": "I've noticed most of the public kernels are using randomly cropped to 5 seconds segments, use them as a training for every epoch and do this with 5 fold. I am wondering would the CV value is credible in that case? I think the data samples for each fold has different trining data for the same label and it may cause the CV value unstable. Maybe caching the cropped segments for each epoch and run the model would give the stable CVs. \nOn the other side in my mind, it is okay with current one because it is random! The expectation of amount of information we get from the training data in each fold would be similar.\nHow do you think?\n\nThank you",
      "votes": null
    },
    {
      "id": "964922",
      "postDate": "08/10/2020 09:13:28",
      "content": "<p>You are absolutely right - you need to crop always the same audio clips for your validation. If you are using pytorch you probably can set validation numpy seed through worker<em>init</em>fn to crop every epoch the same clips.</p>\n<p>But keep in mind that most public validation schemas don't solve the problem of this competition. </p>",
      "rawMarkdown": "You are absolutely right - you need to crop always the same audio clips for your validation. If you are using pytorch you probably can set validation numpy seed through worker_init_fn to crop every epoch the same clips.\n\nBut keep in mind that most public validation schemas don't solve the problem of this competition.",
      "votes": null
    },
    {
      "id": "964957",
      "postDate": "08/10/2020 09:45:34",
      "content": "<p>Oh you are right.</p>\n<blockquote>\n  <p>you need to crop always the same audio clips for your validation.</p>\n</blockquote>\n<p>What should be done here is using cropped \"validation\" samples to have fixed range for each label for every epoch(and of course, every fold).<br>\nHow do you think that random cropping to training samples? In general, training samples is not change for all epochs, but the random cropping does. I am wondering we can see it as one way of augmentation or the one should be prohibited for strict validation.<br>\nIn the other side of my mind again, this wondering may be meaningless, since the competition's evaluation method is too different…</p>\n<p>Thank you for reply</p>",
      "rawMarkdown": "Oh you are right.\n\n&gt; you need to crop always the same audio clips for your validation.\n\nWhat should be done here is using cropped \"validation\" samples to have fixed range for each label for every epoch(and of course, every fold).\nHow do you think that random cropping to training samples? In general, training samples is not change for all epochs, but the random cropping does. I am wondering we can see it as one way of augmentation or the one should be prohibited for strict validation.\nIn the other side of my mind again, this wondering may be meaningless, since the competition's evaluation method is too different...\n\nThank you for reply",
      "votes": null
    },
    {
      "id": "966307",
      "postDate": "08/11/2020 10:15:15",
      "content": "<p>Now I noticed there is <code>np.random.seed</code> function in the most public kernels and checked it makes validation set have the same range for every epoch. No worries of fluctuating validation.</p>",
      "rawMarkdown": "Now I noticed there is `np.random.seed` function in the most public kernels and checked it makes validation set have the same range for every epoch. No worries of fluctuating validation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 964922,
      "author_name": "yaroshevskiy",
      "author_url": "",
      "post_date": "08/10/2020 09:13:28",
      "content": "<p>You are absolutely right - you need to crop always the same audio clips for your validation. If you are using pytorch you probably can set validation numpy seed through worker<em>init</em>fn to crop every epoch the same clips.</p>\n<p>But keep in mind that most public validation schemas don't solve the problem of this competition. </p>",
      "votes": null,
      "replies": [
        {
          "id": 964957,
          "author_name": "jihunlorenzopark",
          "author_url": "",
          "post_date": "08/10/2020 09:45:34",
          "content": "<p>Oh you are right.</p>\n<blockquote>\n  <p>you need to crop always the same audio clips for your validation.</p>\n</blockquote>\n<p>What should be done here is using cropped \"validation\" samples to have fixed range for each label for every epoch(and of course, every fold).<br>\nHow do you think that random cropping to training samples? In general, training samples is not change for all epochs, but the random cropping does. I am wondering we can see it as one way of augmentation or the one should be prohibited for strict validation.<br>\nIn the other side of my mind again, this wondering may be meaningless, since the competition's evaluation method is too different…</p>\n<p>Thank you for reply</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 966307,
      "author_name": "jihunlorenzopark",
      "author_url": "",
      "post_date": "08/11/2020 10:15:15",
      "content": "<p>Now I noticed there is <code>np.random.seed</code> function in the most public kernels and checked it makes validation set have the same range for every epoch. No worries of fluctuating validation.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "964464": "I've noticed most of the public kernels are using randomly cropped to 5 seconds segments, use them as a training for every epoch and do this with 5 fold. I am wondering would the CV value is credible in that case? I think the data samples for each fold has different trining data for the same label and it may cause the CV value unstable. Maybe caching the cropped segments for each epoch and run the model would give the stable CVs. \nOn the other side in my mind, it is okay with current one because it is random! The expectation of amount of information we get from the training data in each fold would be similar.\nHow do you think?\n\nThank you",
    "964922": "You are absolutely right - you need to crop always the same audio clips for your validation. If you are using pytorch you probably can set validation numpy seed through worker_init_fn to crop every epoch the same clips.\n\nBut keep in mind that most public validation schemas don't solve the problem of this competition.",
    "964957": "Oh you are right.\n\n&gt; you need to crop always the same audio clips for your validation.\n\nWhat should be done here is using cropped \"validation\" samples to have fixed range for each label for every epoch(and of course, every fold).\nHow do you think that random cropping to training samples? In general, training samples is not change for all epochs, but the random cropping does. I am wondering we can see it as one way of augmentation or the one should be prohibited for strict validation.\nIn the other side of my mind again, this wondering may be meaningless, since the competition's evaluation method is too different...\n\nThank you for reply",
    "966307": "Now I noticed there is `np.random.seed` function in the most public kernels and checked it makes validation set have the same range for every epoch. No worries of fluctuating validation."
  },
  "source": "meta"
}