{
  "id": 204651,
  "title": "Any Stable Validation Strategy?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/204651",
  "author_name": "Kerem Turgutlu",
  "post_date": "2020-12-16T06:58:48.745000",
  "votes": 13,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I have been trying out a lot of different approaches for this competition either ideas that come to my mind or the ones that are shared in discussions thanks to everyone sharing.</p>\n<p>I used 10 fold random cv for validating my ideas but since it was taking too much time I switched back to 5 fold. Unfortunately, I am not able to find a stable cv strategy that aligns with public LB yet.</p>\n<p>There might be several and more things playing role in this: </p>\n<p>1) Noise in the training set, noise in validation set, noise in public test set - are they random? If yes random validation improvements should reflect to LB. But in my case a model scores 0.910 in CV then score 0.895 in LB and oppositely another model can score 0.895 in CV and 0.901 in LB.</p>\n<p>2) Variance. It is mentioned in papers robust training is needed in the presence of noise and that might be why a lot of TTA rounds actually help, additionally to blending different models. Some improvements can be lower then the effect of the variance. Here we see that a lot top teams are only 0.001 away from each other which is 16 images approximately. I am not able to tell even if my model is statistically significant or it’s by chance. </p>\n<p>I am wondering if anyone manage to come up with a solid working validation strategy. So far only thing I can think of is doing kfold multiple times to reduce the variance as much a possible.</p>",
  "messages": [
    {
      "id": 1115291,
      "postDate": "2020-12-16T06:58:48.747Z",
      "content": "<p>I have been trying out a lot of different approaches for this competition either ideas that come to my mind or the ones that are shared in discussions thanks to everyone sharing.</p>\n<p>I used 10 fold random cv for validating my ideas but since it was taking too much time I switched back to 5 fold. Unfortunately, I am not able to find a stable cv strategy that aligns with public LB yet.</p>\n<p>There might be several and more things playing role in this: </p>\n<p>1) Noise in the training set, noise in validation set, noise in public test set - are they random? If yes random validation improvements should reflect to LB. But in my case a model scores 0.910 in CV then score 0.895 in LB and oppositely another model can score 0.895 in CV and 0.901 in LB.</p>\n<p>2) Variance. It is mentioned in papers robust training is needed in the presence of noise and that might be why a lot of TTA rounds actually help, additionally to blending different models. Some improvements can be lower then the effect of the variance. Here we see that a lot top teams are only 0.001 away from each other which is 16 images approximately. I am not able to tell even if my model is statistically significant or it’s by chance. </p>\n<p>I am wondering if anyone manage to come up with a solid working validation strategy. So far only thing I can think of is doing kfold multiple times to reduce the variance as much a possible.</p>",
      "rawMarkdown": "I have been trying out a lot of different approaches for this competition either ideas that come to my mind or the ones that are shared in discussions thanks to everyone sharing.\n\nI used 10 fold random cv for validating my ideas but since it was taking too much time I switched back to 5 fold. Unfortunately, I am not able to find a stable cv strategy that aligns with public LB yet.\n\nThere might be several and more things playing role in this: \n\n1) Noise in the training set, noise in validation set, noise in public test set - are they random? If yes random validation improvements should reflect to LB. But in my case a model scores 0.910 in CV then score 0.895 in LB and oppositely another model can score 0.895 in CV and 0.901 in LB.\n\n2) Variance. It is mentioned in papers robust training is needed in the presence of noise and that might be why a lot of TTA rounds actually help, additionally to blending different models. Some improvements can be lower then the effect of the variance. Here we see that a lot top teams are only 0.001 away from each other which is 16 images approximately. I am not able to tell even if my model is statistically significant or it’s by chance. \n\nI am wondering if anyone manage to come up with a solid working validation strategy. So far only thing I can think of is doing kfold multiple times to reduce the variance as much a possible.\n\n\n\n\n\n\n\n\n\n\n\n\n\n",
      "votes": 13
    },
    {
      "id": 1115548,
      "postDate": "2020-12-16T11:05:09.760Z",
      "content": "<p>If the noise is completely random means that I hasn't got the same distribution over the wrong labels in training vs public set. So if you select your model based on the validation set metric, that doesn't mean that the performance on the leaderboard set will be correlated with it</p>\n<p>The techniques I am using for benchmark a model configuration are:</p>\n<ol>\n<li>using different seeds numbers to generate the data fold splits and then average the results ( you will have different noise distribution on valid but averaging them, it gives you a more reliable result)</li>\n<li>eliminating noise samples from data (training on correct data and predicting on noisy is better then training on noisy and predicting on noisy)</li>\n</ol>",
      "rawMarkdown": "If the noise is completely random means that I hasn't got the same distribution over the wrong labels in training vs public set. So if you select your model based on the validation set metric, that doesn't mean that the performance on the leaderboard set will be correlated with it\n\nThe techniques I am using for benchmark a model configuration are:\n1. using different seeds numbers to generate the data fold splits and then average the results ( you will have different noise distribution on valid but averaging them, it gives you a more reliable result)\n2. eliminating noise samples from data (training on correct data and predicting on noisy is better then training on noisy and predicting on noisy)\n\n",
      "votes": 5,
      "replies": [
        {
          "id": 1116084,
          "postDate": "2020-12-16T20:26:35.143Z",
          "content": "<p>I agree with both of your points. Robust training + bootstrapping and reducing variance via multiple kfold should be what we need to trust instead of LB.</p>",
          "rawMarkdown": "I agree with both of your points. Robust training + bootstrapping and reducing variance via multiple kfold should be what we need to trust instead of LB.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1116755,
      "postDate": "2020-12-17T12:52:09.197Z",
      "content": "<p>1) Is \"random cv\" StratifiedK Split? My CV and LB difference is quite a lot CV: 0.8905 LB: 0.872. Any comments?<br>\n2) How do you know there exists a noise in data that is causing issues in correlation btw CV and LB? Maybe the test data is from different city or the split is Time based?</p>",
      "rawMarkdown": "1) Is \"random cv\" StratifiedK Split? My CV and LB difference is quite a lot CV: 0.8905 LB: 0.872. Any comments?\n2) How do you know there exists a noise in data that is causing issues in correlation btw CV and LB? Maybe the test data is from different city or the split is Time based?",
      "votes": 1,
      "replies": [
        {
          "id": 1118274,
          "postDate": "2020-12-18T22:33:38.337Z",
          "content": "<p>1) Random cv is not stratified, just random kfold. </p>\n<p>2) I am not sure, it’s just an hypothesis but maybe noise distributions can be different in training and test images. </p>",
          "rawMarkdown": "1) Random cv is not stratified, just random kfold. \n\n2) I am not sure, it’s just an hypothesis but maybe noise distributions can be different in training and test images. ",
          "votes": 2
        },
        {
          "id": 1118399,
          "postDate": "2020-12-19T03:21:36.893Z",
          "content": "<p>My opinion is no way for 1st to get so bad score on just 5 class classification problem unless noisy dataset.</p>",
          "rawMarkdown": "My opinion is no way for 1st to get so bad score on just 5 class classification problem unless noisy dataset."
        }
      ]
    },
    {
      "id": 1116197,
      "postDate": "2020-12-17T00:22:31.160Z",
      "content": "<p>Given the small number of experiments we can use a t-distribution for estimating the confidence interval of validation accuracy might be a good idea. That can give a sense of possible shake-up: </p>\n<pre><code>confidence_level = 0.95\ndegrees_freedom = sample.size - 1 # sample is an array of validation scores from multiple experiments\nsample_mean = np.mean(sample)\nsample_standard_error = scipy.stats.sem(sample)\n\nconfidence_interval = scipy.stats.t.interval(confidence_level, degrees_freedom, sample_mean, sample_standard_error)\n</code></pre>",
      "rawMarkdown": "Given the small number of experiments we can use a t-distribution for estimating the confidence interval of validation accuracy might be a good idea. That can give a sense of possible shake-up: \n```\nconfidence_level = 0.95\ndegrees_freedom = sample.size - 1 # sample is an array of validation scores from multiple experiments\nsample_mean = np.mean(sample)\nsample_standard_error = scipy.stats.sem(sample)\n\nconfidence_interval = scipy.stats.t.interval(confidence_level, degrees_freedom, sample_mean, sample_standard_error)\n``` ",
      "votes": 1,
      "replies": [
        {
          "id": 1116211,
          "postDate": "2020-12-17T00:45:17.523Z",
          "content": "<p>Not really necessary to use any approximation. See below, you can use the beta distribution to just get the exact Clopper-Pearson CI as Beta(0.025, correct, N-correct+1) to Beta(0.975, correct+1, N-correct). Or some other variant like adding 0.5 such as Beta(0.025, correct+0.5, N-correct+0.5) to Beta(0.975, correct+0.5, N-correct+0.5) - although that will not really make much of a difference.</p>",
          "rawMarkdown": "Not really necessary to use any approximation. See below, you can use the beta distribution to just get the exact Clopper-Pearson CI as Beta(0.025, correct, N-correct+1) to Beta(0.975, correct+1, N-correct). Or some other variant like adding 0.5 such as Beta(0.025, correct+0.5, N-correct+0.5) to Beta(0.975, correct+0.5, N-correct+0.5) - although that will not really make much of a difference.",
          "votes": 1
        },
        {
          "id": 1116219,
          "postDate": "2020-12-17T01:04:02.690Z",
          "content": "<p>Thanks for pointing it out, I am not much familiar with using beta distribution for CI :) Do you mean with this method we don't even need to run multiple experiments to get a list sample of scores, then to sample mean and stderr calcs? All we need is # of correct predictions and sample size from just 1 experiment?</p>",
          "rawMarkdown": "Thanks for pointing it out, I am not much familiar with using beta distribution for CI :) Do you mean with this method we don't even need to run multiple experiments to get a list sample of scores, then to sample mean and stderr calcs? All we need is # of correct predictions and sample size from just 1 experiment?",
          "votes": 1
        },
        {
          "id": 1116815,
          "postDate": "2020-12-17T13:40:31.607Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1116814,
          "postDate": "2020-12-17T13:40:31.607Z",
          "content": "<p>Ooops, I misread your comment somewhat. What I meant more is that you don't need to use a normal approximation to get a confidence interval based on independent images either being correctly or incorrectly classified (i.e. for a single experiment). Once you get to multiple experiments, it actually gets a bit messier, because whether the same image gets correctly classified for different seeds is not independent. If I understand what you did, you e.g. had 6 seeds, so got 6 estimates of accuracy (let's say [0.85, 0.86, 0.84, 0.85, 0.86, 0.85]) and then you look at the mean and variance around that. I guess that kind of will capture the correlation in a sense and is pretty easy to implement, but e.g. ignores how uncertain you are about each of the estimates. I assume it would likely be more efficient (=more information with fewer seeds), if one fit a random effects logistic regression (with a random effect on the intercept for each image and perhaps class as an explanatory variable).</p>",
          "rawMarkdown": "Ooops, I misread your comment somewhat. What I meant more is that you don't need to use a normal approximation to get a confidence interval based on independent images either being correctly or incorrectly classified (i.e. for a single experiment). Once you get to multiple experiments, it actually gets a bit messier, because whether the same image gets correctly classified for different seeds is not independent. If I understand what you did, you e.g. had 6 seeds, so got 6 estimates of accuracy (let's say [0.85, 0.86, 0.84, 0.85, 0.86, 0.85]) and then you look at the mean and variance around that. I guess that kind of will capture the correlation in a sense and is pretty easy to implement, but e.g. ignores how uncertain you are about each of the estimates. I assume it would likely be more efficient (=more information with fewer seeds), if one fit a random effects logistic regression (with a random effect on the intercept for each image and perhaps class as an explanatory variable).",
          "votes": 1
        }
      ]
    },
    {
      "id": 1115826,
      "postDate": "2020-12-16T15:47:55.267Z",
      "content": "<p>You get an approximate idea of uncertainty around accuracy by looking at a credible interval (or a confidence interval) for the accuracy. E.g. for 0.895 accuracy, a 80% CrI goes from 0.892 to 0.898 (while a 99.9% CrI has an upper end at 0.902).</p>\n<pre><code>from scipy.stats import beta\nlcrl = np.round( beta.ppf(0.1, 19150, 2247), 5)\nmedian = np.round( beta.ppf(0.5, 19150, 2247), 5)\nucrl = np.round( beta.ppf(0.9, 19150, 2247), 5)\nprint(f'Median {median} with lower 80% CrI limit {lcrl} to upper 80% CrI limit {ucrl}')\n</code></pre>\n<p>So, variation that goes from 0.895 to 0.901 is perhaps a bit surprising, but not completely beyond what you might get through things that are just randomness. Of course, there might be some systematic differences - e.g. the organizers might have split older vs. newer photos.</p>\n<p>Repeated 5-fold (e.g. 2 times or 3 times 5-fold - as mentioned by <a href=\"https://www.kaggle.com/vladvdv\" target=\"_blank\">@vladvdv</a>) could be an option to reduce the variation (but because there would be a correlation between the out-of-fold validation scores, its a bit harder to say how much that reduces your uncertainty around performance). Of course, that has the same downside as 10 fold (i.e. having to train 10 or 15 times), but I would actually expect it to improve CV to LB correlation more than using a single 10-fold CV scheme.</p>\n<p>Stratified (by class) k-fold (if you are not using it) could be another option (or perhaps <a href=\"https://www.kaggle.com/bjoernholzhauer/cassava-leaf-disease-classif-eda-cv-strategy\" target=\"_blank\">stratified by image type</a>).</p>",
      "rawMarkdown": "You get an approximate idea of uncertainty around accuracy by looking at a credible interval (or a confidence interval) for the accuracy. E.g. for 0.895 accuracy, a 80% CrI goes from 0.892 to 0.898 (while a 99.9% CrI has an upper end at 0.902).\n```\nfrom scipy.stats import beta\nlcrl = np.round( beta.ppf(0.1, 19150, 2247), 5)\nmedian = np.round( beta.ppf(0.5, 19150, 2247), 5)\nucrl = np.round( beta.ppf(0.9, 19150, 2247), 5)\nprint(f'Median {median} with lower 80% CrI limit {lcrl} to upper 80% CrI limit {ucrl}')\n\n```\nSo, variation that goes from 0.895 to 0.901 is perhaps a bit surprising, but not completely beyond what you might get through things that are just randomness. Of course, there might be some systematic differences - e.g. the organizers might have split older vs. newer photos.\n\nRepeated 5-fold (e.g. 2 times or 3 times 5-fold - as mentioned by @vladvdv) could be an option to reduce the variation (but because there would be a correlation between the out-of-fold validation scores, its a bit harder to say how much that reduces your uncertainty around performance). Of course, that has the same downside as 10 fold (i.e. having to train 10 or 15 times), but I would actually expect it to improve CV to LB correlation more than using a single 10-fold CV scheme.\n\nStratified (by class) k-fold (if you are not using it) could be another option (or perhaps [stratified by image type](https://www.kaggle.com/bjoernholzhauer/cassava-leaf-disease-classif-eda-cv-strategy)).",
      "votes": 1,
      "replies": [
        {
          "id": 1116222,
          "postDate": "2020-12-17T01:05:32.160Z",
          "content": "<p>Nice kernel thanks for linking it!</p>",
          "rawMarkdown": "Nice kernel thanks for linking it!"
        }
      ]
    },
    {
      "id": 1118907,
      "postDate": "2020-12-19T14:13:11.300Z",
      "content": "<p>Hello!<br>\nIs single fold 901?</p>",
      "rawMarkdown": "Hello!\nIs single fold 901?"
    }
  ],
  "comments": [
    {
      "id": 1115548,
      "author_name": "Vlad Vaduva",
      "author_url": "",
      "post_date": "2020-12-16T11:05:09.760000",
      "content": "<p>If the noise is completely random means that I hasn't got the same distribution over the wrong labels in training vs public set. So if you select your model based on the validation set metric, that doesn't mean that the performance on the leaderboard set will be correlated with it</p>\n<p>The techniques I am using for benchmark a model configuration are:</p>\n<ol>\n<li>using different seeds numbers to generate the data fold splits and then average the results ( you will have different noise distribution on valid but averaging them, it gives you a more reliable result)</li>\n<li>eliminating noise samples from data (training on correct data and predicting on noisy is better then training on noisy and predicting on noisy)</li>\n</ol>",
      "votes": 5,
      "replies": [
        {
          "id": 1116084,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2020-12-16T20:26:35.143000",
          "content": "<p>I agree with both of your points. Robust training + bootstrapping and reducing variance via multiple kfold should be what we need to trust instead of LB.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1116755,
      "author_name": "ammad",
      "author_url": "",
      "post_date": "2020-12-17T12:52:09.197000",
      "content": "<p>1) Is \"random cv\" StratifiedK Split? My CV and LB difference is quite a lot CV: 0.8905 LB: 0.872. Any comments?<br>\n2) How do you know there exists a noise in data that is causing issues in correlation btw CV and LB? Maybe the test data is from different city or the split is Time based?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1118274,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2020-12-18T22:33:38.337000",
          "content": "<p>1) Random cv is not stratified, just random kfold. </p>\n<p>2) I am not sure, it’s just an hypothesis but maybe noise distributions can be different in training and test images. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1118399,
          "author_name": "Yi Wu",
          "author_url": "",
          "post_date": "2020-12-19T03:21:36.893000",
          "content": "<p>My opinion is no way for 1st to get so bad score on just 5 class classification problem unless noisy dataset.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1116197,
      "author_name": "Kerem Turgutlu",
      "author_url": "",
      "post_date": "2020-12-17T00:22:31.160000",
      "content": "<p>Given the small number of experiments we can use a t-distribution for estimating the confidence interval of validation accuracy might be a good idea. That can give a sense of possible shake-up: </p>\n<pre><code>confidence_level = 0.95\ndegrees_freedom = sample.size - 1 # sample is an array of validation scores from multiple experiments\nsample_mean = np.mean(sample)\nsample_standard_error = scipy.stats.sem(sample)\n\nconfidence_interval = scipy.stats.t.interval(confidence_level, degrees_freedom, sample_mean, sample_standard_error)\n</code></pre>",
      "votes": 1,
      "replies": [
        {
          "id": 1116211,
          "author_name": "Björn",
          "author_url": "",
          "post_date": "2020-12-17T00:45:17.523000",
          "content": "<p>Not really necessary to use any approximation. See below, you can use the beta distribution to just get the exact Clopper-Pearson CI as Beta(0.025, correct, N-correct+1) to Beta(0.975, correct+1, N-correct). Or some other variant like adding 0.5 such as Beta(0.025, correct+0.5, N-correct+0.5) to Beta(0.975, correct+0.5, N-correct+0.5) - although that will not really make much of a difference.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1116219,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2020-12-17T01:04:02.690000",
          "content": "<p>Thanks for pointing it out, I am not much familiar with using beta distribution for CI :) Do you mean with this method we don't even need to run multiple experiments to get a list sample of scores, then to sample mean and stderr calcs? All we need is # of correct predictions and sample size from just 1 experiment?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1116815,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-17T13:40:31.607000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1116814,
          "author_name": "Björn",
          "author_url": "",
          "post_date": "2020-12-17T13:40:31.607000",
          "content": "<p>Ooops, I misread your comment somewhat. What I meant more is that you don't need to use a normal approximation to get a confidence interval based on independent images either being correctly or incorrectly classified (i.e. for a single experiment). Once you get to multiple experiments, it actually gets a bit messier, because whether the same image gets correctly classified for different seeds is not independent. If I understand what you did, you e.g. had 6 seeds, so got 6 estimates of accuracy (let's say [0.85, 0.86, 0.84, 0.85, 0.86, 0.85]) and then you look at the mean and variance around that. I guess that kind of will capture the correlation in a sense and is pretty easy to implement, but e.g. ignores how uncertain you are about each of the estimates. I assume it would likely be more efficient (=more information with fewer seeds), if one fit a random effects logistic regression (with a random effect on the intercept for each image and perhaps class as an explanatory variable).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1115826,
      "author_name": "Björn",
      "author_url": "",
      "post_date": "2020-12-16T15:47:55.267000",
      "content": "<p>You get an approximate idea of uncertainty around accuracy by looking at a credible interval (or a confidence interval) for the accuracy. E.g. for 0.895 accuracy, a 80% CrI goes from 0.892 to 0.898 (while a 99.9% CrI has an upper end at 0.902).</p>\n<pre><code>from scipy.stats import beta\nlcrl = np.round( beta.ppf(0.1, 19150, 2247), 5)\nmedian = np.round( beta.ppf(0.5, 19150, 2247), 5)\nucrl = np.round( beta.ppf(0.9, 19150, 2247), 5)\nprint(f'Median {median} with lower 80% CrI limit {lcrl} to upper 80% CrI limit {ucrl}')\n</code></pre>\n<p>So, variation that goes from 0.895 to 0.901 is perhaps a bit surprising, but not completely beyond what you might get through things that are just randomness. Of course, there might be some systematic differences - e.g. the organizers might have split older vs. newer photos.</p>\n<p>Repeated 5-fold (e.g. 2 times or 3 times 5-fold - as mentioned by <a href=\"https://www.kaggle.com/vladvdv\" target=\"_blank\">@vladvdv</a>) could be an option to reduce the variation (but because there would be a correlation between the out-of-fold validation scores, its a bit harder to say how much that reduces your uncertainty around performance). Of course, that has the same downside as 10 fold (i.e. having to train 10 or 15 times), but I would actually expect it to improve CV to LB correlation more than using a single 10-fold CV scheme.</p>\n<p>Stratified (by class) k-fold (if you are not using it) could be another option (or perhaps <a href=\"https://www.kaggle.com/bjoernholzhauer/cassava-leaf-disease-classif-eda-cv-strategy\" target=\"_blank\">stratified by image type</a>).</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1116222,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2020-12-17T01:05:32.160000",
          "content": "<p>Nice kernel thanks for linking it!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1118907,
      "author_name": "HAaHAa",
      "author_url": "",
      "post_date": "2020-12-19T14:13:11.300000",
      "content": "<p>Hello!<br>\nIs single fold 901?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1115291": "I have been trying out a lot of different approaches for this competition either ideas that come to my mind or the ones that are shared in discussions thanks to everyone sharing.\n\nI used 10 fold random cv for validating my ideas but since it was taking too much time I switched back to 5 fold. Unfortunately, I am not able to find a stable cv strategy that aligns with public LB yet.\n\nThere might be several and more things playing role in this: \n\n1) Noise in the training set, noise in validation set, noise in public test set - are they random? If yes random validation improvements should reflect to LB. But in my case a model scores 0.910 in CV then score 0.895 in LB and oppositely another model can score 0.895 in CV and 0.901 in LB.\n\n2) Variance. It is mentioned in papers robust training is needed in the presence of noise and that might be why a lot of TTA rounds actually help, additionally to blending different models. Some improvements can be lower then the effect of the variance. Here we see that a lot top teams are only 0.001 away from each other which is 16 images approximately. I am not able to tell even if my model is statistically significant or it’s by chance. \n\nI am wondering if anyone manage to come up with a solid working validation strategy. So far only thing I can think of is doing kfold multiple times to reduce the variance as much a possible.\n\n\n\n\n\n\n\n\n\n\n\n\n\n",
    "1115548": "If the noise is completely random means that I hasn't got the same distribution over the wrong labels in training vs public set. So if you select your model based on the validation set metric, that doesn't mean that the performance on the leaderboard set will be correlated with it\n\nThe techniques I am using for benchmark a model configuration are:\n1. using different seeds numbers to generate the data fold splits and then average the results ( you will have different noise distribution on valid but averaging them, it gives you a more reliable result)\n2. eliminating noise samples from data (training on correct data and predicting on noisy is better then training on noisy and predicting on noisy)\n\n",
    "1116755": "1) Is \"random cv\" StratifiedK Split? My CV and LB difference is quite a lot CV: 0.8905 LB: 0.872. Any comments?\n2) How do you know there exists a noise in data that is causing issues in correlation btw CV and LB? Maybe the test data is from different city or the split is Time based?",
    "1116197": "Given the small number of experiments we can use a t-distribution for estimating the confidence interval of validation accuracy might be a good idea. That can give a sense of possible shake-up: \n```\nconfidence_level = 0.95\ndegrees_freedom = sample.size - 1 # sample is an array of validation scores from multiple experiments\nsample_mean = np.mean(sample)\nsample_standard_error = scipy.stats.sem(sample)\n\nconfidence_interval = scipy.stats.t.interval(confidence_level, degrees_freedom, sample_mean, sample_standard_error)\n``` ",
    "1115826": "You get an approximate idea of uncertainty around accuracy by looking at a credible interval (or a confidence interval) for the accuracy. E.g. for 0.895 accuracy, a 80% CrI goes from 0.892 to 0.898 (while a 99.9% CrI has an upper end at 0.902).\n```\nfrom scipy.stats import beta\nlcrl = np.round( beta.ppf(0.1, 19150, 2247), 5)\nmedian = np.round( beta.ppf(0.5, 19150, 2247), 5)\nucrl = np.round( beta.ppf(0.9, 19150, 2247), 5)\nprint(f'Median {median} with lower 80% CrI limit {lcrl} to upper 80% CrI limit {ucrl}')\n\n```\nSo, variation that goes from 0.895 to 0.901 is perhaps a bit surprising, but not completely beyond what you might get through things that are just randomness. Of course, there might be some systematic differences - e.g. the organizers might have split older vs. newer photos.\n\nRepeated 5-fold (e.g. 2 times or 3 times 5-fold - as mentioned by @vladvdv) could be an option to reduce the variation (but because there would be a correlation between the out-of-fold validation scores, its a bit harder to say how much that reduces your uncertainty around performance). Of course, that has the same downside as 10 fold (i.e. having to train 10 or 15 times), but I would actually expect it to improve CV to LB correlation more than using a single 10-fold CV scheme.\n\nStratified (by class) k-fold (if you are not using it) could be another option (or perhaps [stratified by image type](https://www.kaggle.com/bjoernholzhauer/cassava-leaf-disease-classif-eda-cv-strategy)).",
    "1118907": "Hello!\nIs single fold 901?"
  }
}