{
  "id": 275233,
  "title": "Leak in metadata?",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/275233",
  "author_name": "",
  "post_date": "2021-09-29T13:33:35.859104Z",
  "votes": 18,
  "comment_count": 17,
  "views": 0,
  "content": "<p>I found out that the metadata is surprisingly good for training a classifier. Since I highly doubt that the use of this data is allowed, I decided to make this information public.</p>\n<p>The MRI parameters correlate with MGMT promoter methylation. This means that the people who configured the MRI scanner or preprocessed the data leaked some information into the data. I found three parameters that are good indicators of MGMT.</p>\n<p>(1) The echo train length (ETL) determines the number of echoes within a repetition time. An increase of ETL introduces more T2 decay in the image and decreases the scan time. </p>\n<p><img src=\"https://i.imgur.com/cyAloEV.jpeg\" alt=\"ETL\"></p>\n<p>(2) The phase field of view can also be used to reduce scan time. </p>\n<p><img src=\"https://i.imgur.com/LAbzKnJ.jpg\" alt=\"Phase field of view\"></p>\n<p>(3) Number of scans in folders.</p>\n<p>I trained a logistic regression and got a CV of 0.6575 and LB of 0.641. LB and CV are perfectly correlated. However, I do not know if private data is also affected.</p>\n<p>The notebook can be found <a href=\"https://www.kaggle.com/lars123/leak-in-metadata\" target=\"_blank\">here</a>.</p>",
  "messages": [
    {
      "id": "1528225",
      "postDate": "09/29/2021 13:33:35",
      "content": "<p>I found out that the metadata is surprisingly good for training a classifier. Since I highly doubt that the use of this data is allowed, I decided to make this information public.</p>\n<p>The MRI parameters correlate with MGMT promoter methylation. This means that the people who configured the MRI scanner or preprocessed the data leaked some information into the data. I found three parameters that are good indicators of MGMT.</p>\n<p>(1) The echo train length (ETL) determines the number of echoes within a repetition time. An increase of ETL introduces more T2 decay in the image and decreases the scan time. </p>\n<p><img src=\"https://i.imgur.com/cyAloEV.jpeg\" alt=\"ETL\"></p>\n<p>(2) The phase field of view can also be used to reduce scan time. </p>\n<p><img src=\"https://i.imgur.com/LAbzKnJ.jpg\" alt=\"Phase field of view\"></p>\n<p>(3) Number of scans in folders.</p>\n<p>I trained a logistic regression and got a CV of 0.6575 and LB of 0.641. LB and CV are perfectly correlated. However, I do not know if private data is also affected.</p>\n<p>The notebook can be found <a href=\"https://www.kaggle.com/lars123/leak-in-metadata\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "I found out that the metadata is surprisingly good for training a classifier. Since I highly doubt that the use of this data is allowed, I decided to make this information public.\n\nThe MRI parameters correlate with MGMT promoter methylation. This means that the people who configured the MRI scanner or preprocessed the data leaked some information into the data. I found three parameters that are good indicators of MGMT.\n\n(1) The echo train length (ETL) determines the number of echoes within a repetition time. An increase of ETL introduces more T2 decay in the image and decreases the scan time. \n\n![ETL](https://i.imgur.com/cyAloEV.jpeg)\n\n(2) The phase field of view can also be used to reduce scan time. \n\n![Phase field of view](https://i.imgur.com/LAbzKnJ.jpg)\n\n(3) Number of scans in folders.\n\nI trained a logistic regression and got a CV of 0.6575 and LB of 0.641. LB and CV are perfectly correlated. However, I do not know if private data is also affected.\n\nThe notebook can be found [here](https://www.kaggle.com/lars123/leak-in-metadata).",
      "votes": null
    },
    {
      "id": "1528363",
      "postDate": "09/29/2021 15:28:07",
      "content": "<p>I reran your notebook with number of folds changed to 5, got 0.57 validation AUC while public score is still around 0.65. Not perfectly correlated as you suggested but I agree that maybe those metadata features are useful.</p>",
      "rawMarkdown": "I reran your notebook with number of folds changed to 5, got 0.57 validation AUC while public score is still around 0.65. Not perfectly correlated as you suggested but I agree that maybe those metadata features are useful.",
      "votes": null
    },
    {
      "id": "1528507",
      "postDate": "09/29/2021 17:42:05",
      "content": "<p>The std of AUC was very high (0.4+) in the original notebook with 200-fold CV, so I ran it too with 5-fold CV. The average validation AUC was 0.57 and std remained quite high (0.095).</p>\n<p>There could be some correlation between the metadata features and MGMT label but the features don't seem very robust considering the high std of the AUC between folds.</p>\n<p>This is an interesting discovery, and I wonder if CNNs are able to pick up some of the imaging parameter features like ETL and overfit to it.</p>",
      "rawMarkdown": "The std of AUC was very high (0.4+) in the original notebook with 200-fold CV, so I ran it too with 5-fold CV. The average validation AUC was 0.57 and std remained quite high (0.095).\n\nThere could be some correlation between the metadata features and MGMT label but the features don't seem very robust considering the high std of the AUC between folds.\n\nThis is an interesting discovery, and I wonder if CNNs are able to pick up some of the imaging parameter features like ETL and overfit to it.",
      "votes": null
    },
    {
      "id": "1528588",
      "postDate": "09/29/2021 19:06:30",
      "content": "<p>I don't think theses DICOM values or number of scans(images) have a direct correlation to MGMT for a couple of reasons.  If they do, I suspect it's purely coincidental due to the small sample size.  </p>\n<p>The technologist performing the scans likely did not know the patient's MGMT status (since that usually comes from a post-mortem biopsy). The scan parameters are generally programmed into a protocol by an engineer and not tweaked by the technologist at acquisition time. Protocols differ slightly between machines, companies and radiologists. </p>\n<p>FOV is generally expected to be square for a brain. But if the patient was having a neck and brain scan at the same time, the chosen protocol might use a rectangular FOV. A dedicated brain scan should always use a square FOV.</p>",
      "rawMarkdown": "I don't think theses DICOM values or number of scans(images) have a direct correlation to MGMT for a couple of reasons.  If they do, I suspect it's purely coincidental due to the small sample size.  \n\nThe technologist performing the scans likely did not know the patient's MGMT status (since that usually comes from a post-mortem biopsy). The scan parameters are generally programmed into a protocol by an engineer and not tweaked by the technologist at acquisition time. Protocols differ slightly between machines, companies and radiologists. \n\nFOV is generally expected to be square for a brain. But if the patient was having a neck and brain scan at the same time, the chosen protocol might use a rectangular FOV. A dedicated brain scan should always use a square FOV.",
      "votes": null
    },
    {
      "id": "1528591",
      "postDate": "09/29/2021 19:16:38",
      "content": "<p>I found that one can get to ~59% accuracy with a simple sklearn SVM classifier <em>using the number of images alone</em>, patients with MGMT = 1 tend to have substantially more FLAIR and T2w images.  The number of images was determined by the doctor/radiologist presumably <em>before</em> the methylation status of the tumor was known, therefore the number of slices should be completely decoupled from MGMT status.  This is why it's important to have some understanding of what your data actually is, otherwise spurious correlations can enter.</p>",
      "rawMarkdown": "I found that one can get to ~59% accuracy with a simple sklearn SVM classifier *using the number of images alone*, patients with MGMT = 1 tend to have substantially more FLAIR and T2w images.  The number of images was determined by the doctor/radiologist presumably *before* the methylation status of the tumor was known, therefore the number of slices should be completely decoupled from MGMT status.  This is why it's important to have some understanding of what your data actually is, otherwise spurious correlations can enter.",
      "votes": null
    },
    {
      "id": "1528592",
      "postDate": "09/29/2021 19:18:17",
      "content": "<p>The way you are computing your CV AUC is incorrect, you need to do something like this : </p>\n<pre><code>pred_oof = np.zeros(X.shape[0])\nfor fold, (train_index, val_index) in enumerate(StratifiedKFold(n_splits=k).split(X, y)):\n    [...]\n\n    y_pred = regr.predict_proba(X_val)[...,1]    \n    pred_oof[val_index] = y_pred   # save predictions\n\nprint(\"Loss :\", log_loss(y, pred_oof))\nprint(\"AUC :\", roc_auc_score(y, pred_oof))   # compute the AUC on the whole samples\n</code></pre>\n<p>Which results in 0.55 CV and probably means the 0.6+ lb is just noise.</p>",
      "rawMarkdown": "The way you are computing your CV AUC is incorrect, you need to do something like this : \n\n```\npred_oof = np.zeros(X.shape[0])\nfor fold, (train_index, val_index) in enumerate(StratifiedKFold(n_splits=k).split(X, y)):\n    [...]\n    \n    y_pred = regr.predict_proba(X_val)[...,1]    \n    pred_oof[val_index] = y_pred   # save predictions\n\nprint(\"Loss :\", log_loss(y, pred_oof))\nprint(\"AUC :\", roc_auc_score(y, pred_oof))   # compute the AUC on the whole samples\n```\n\n\nWhich results in 0.55 CV and probably means the 0.6+ lb is just noise.",
      "votes": null
    },
    {
      "id": "1528646",
      "postDate": "09/29/2021 20:55:45",
      "content": "<p>One possible explanation for the correlation with imaging protocol: </p>\n<p>If the data came from different sites, some sites might preferentially see more MGMT+ patients, and those sites happen to have the protocols we are noticing. </p>\n<p>So we are separating the data based on scanning site and for some reason different sites have a different mix of patients. </p>",
      "rawMarkdown": "One possible explanation for the correlation with imaging protocol: \n\nIf the data came from different sites, some sites might preferentially see more MGMT+ patients, and those sites happen to have the protocols we are noticing. \n\nSo we are separating the data based on scanning site and for some reason different sites have a different mix of patients.",
      "votes": null
    },
    {
      "id": "1528676",
      "postDate": "09/29/2021 21:46:57",
      "content": "<p>I think the way you would compute it would result into an OOF score and not a CV score in the <a href=\"https://www.jmlr.org/papers/volume5/grandvalet04a/grandvalet04a.pdf\" target=\"_blank\">traditional ML sense</a>. <a href=\"https://scikit-learn.org/stable/modules/cross_validation.html#cross-validation\" target=\"_blank\">Sklearn</a> also computes it in the same way as far as I can see. Is there any advantage to using regular CV?</p>",
      "rawMarkdown": "I think the way you would compute it would result into an OOF score and not a CV score in the [traditional ML sense](https://www.jmlr.org/papers/volume5/grandvalet04a/grandvalet04a.pdf). [Sklearn](https://scikit-learn.org/stable/modules/cross_validation.html#cross-validation) also computes it in the same way as far as I can see. Is there any advantage to using regular CV?",
      "votes": null
    },
    {
      "id": "1528685",
      "postDate": "09/29/2021 21:55:27",
      "content": "<p>By decreasing the number of folds (e.g. K=5), we increase the bias of the CV estimator. As K-&gt;N, there is less bias and higher variance. Hence, the big increase in std for K=200. IMHO here a high K is better as it gives a better estimate of the bias (even if std suffers i.e. bias-variance tradeoff). See <a href=\"https://stats.stackexchange.com/questions/61783/bias-and-variance-in-leave-one-out-vs-k-fold-cross-validation\" target=\"_blank\">https://stats.stackexchange.com/questions/61783/bias-and-variance-in-leave-one-out-vs-k-fold-cross-validation</a> Normally, I would also use K=5 or K=10 but the dataset is so small. This is at least as I would interpret the results.</p>",
      "rawMarkdown": "By decreasing the number of folds (e.g. K=5), we increase the bias of the CV estimator. As K->N, there is less bias and higher variance. Hence, the big increase in std for K=200. IMHO here a high K is better as it gives a better estimate of the bias (even if std suffers i.e. bias-variance tradeoff). See https://stats.stackexchange.com/questions/61783/bias-and-variance-in-leave-one-out-vs-k-fold-cross-validation Normally, I would also use K=5 or K=10 but the dataset is so small. This is at least as I would interpret the results.",
      "votes": null
    },
    {
      "id": "1529144",
      "postDate": "09/30/2021 07:22:53",
      "content": "<p>The oof score is more reliable for metrics such as the AUC (especially since you are using 200 folds) - and I believe it is common practice for most Kaggle competitions. </p>",
      "rawMarkdown": "The oof score is more reliable for metrics such as the AUC (especially since you are using 200 folds) - and I believe it is common practice for most Kaggle competitions.",
      "votes": null
    },
    {
      "id": "1529265",
      "postDate": "09/30/2021 08:58:12",
      "content": "<p>I agree that oof score is common practice for kaggle. But I haven't read anywhere that oof score is more reliable than his method. Actually I think <a href=\"https://www.kaggle.com/lars123\" target=\"_blank\">@lars123</a>'s traditional method is more reliable because it can also calculate the variance.</p>",
      "rawMarkdown": "I agree that oof score is common practice for kaggle. But I haven't read anywhere that oof score is more reliable than his method. Actually I think @lars123's traditional method is more reliable because it can also calculate the variance.",
      "votes": null
    },
    {
      "id": "1529334",
      "postDate": "09/30/2021 10:00:24",
      "content": "<p>The reliability comes from the AUC. For instance using 4 samples and <code>k = 2</code> folds : </p>\n<pre><code>auc_avg = (auc([x1, x2], [y1, y2]) + auc([x3, x4], [y3, y4])) / 2 \nauc_oof = auc([x1, x2, x3, x4], [y1, y2, y3, y4])\n</code></pre>\n<p><br>\nYou don't actually have the following : <code>auc_avg = avg_oof</code></p>\n<p>And in fact (I think - small explaination bellow) you have : <code>auc_avg &gt;= avg_oof</code></p>\n<p>Let's say <code>y1 = 0, y2 = 1, y3 = 0, y4 = 1</code> : <br>\nTo have  <code>auc_oof = 1</code> you need to find <code>x1 &lt;= x3 &lt; x2 &lt;= x4</code><br>\nTo have <code>auc_avg = 1</code> you need to find <code>x1 &lt; x2 and x3 &lt; x4</code> which is much easier.</p>\n<p>In the test set the computation is made on the <code>n</code> samples directly, and is not the average of <code>k</code> <code>n/k</code> AUCs.</p>\n<p>The AUC is a rank based metric, and its stability benefits from being computed on more samples. <br>\nFor instance, <a href=\"https://www.kaggle.com/lars123\" target=\"_blank\">@lars123</a>'s notebook shows a 0.4+ variance for a 0.6+ auc which is huge. <br>\nThis is because he computes 200 aucs on ~3 samples. <br>\nSimply reducing the number of folds as you did reduces the variance a lot, but also the cv score. </p>",
      "rawMarkdown": "The reliability comes from the AUC. For instance using 4 samples and `k = 2` folds : \n```\nauc_avg = (auc([x1, x2], [y1, y2]) + auc([x3, x4], [y3, y4])) / 2 \nauc_oof = auc([x1, x2, x3, x4], [y1, y2, y3, y4])\n``` \nYou don't actually have the following : `auc_avg = avg_oof`\n\nAnd in fact (I think - small explaination bellow) you have : `auc_avg >= avg_oof`\n\nLet's say `y1 = 0, y2 = 1, y3 = 0, y4 = 1` : \nTo have  `auc_oof = 1` you need to find `x1 <= x3 < x2 <= x4`\nTo have `auc_avg = 1` you need to find `x1 < x2 and x3 < x4` which is much easier.\n\nIn the test set the computation is made on the `n` samples directly, and is not the average of `k` `n/k` AUCs.\n\n\nThe AUC is a rank based metric, and its stability benefits from being computed on more samples. \nFor instance, @lars123's notebook shows a 0.4+ variance for a 0.6+ auc which is huge. \nThis is because he computes 200 aucs on ~3 samples. \nSimply reducing the number of folds as you did reduces the variance a lot, but also the cv score.",
      "votes": null
    },
    {
      "id": "1529800",
      "postDate": "09/30/2021 16:57:51",
      "content": "<p>While your argument makes intuitive sense, I'm not sure I totally agree with that. I've just created a new <a href=\"https://www.kaggle.com/lars123/oof-vs-cv\" target=\"_blank\">notebook</a> to show some empirical results.</p>\n<p>I did a train-test-split on the dataset (according to the public-private LB ratio i.e. 78%, 22%). Then I trained a logistic regression with 5-fold and 60-fold CV and looked at the OOF and CV score. The CV score was in all cases closer to the true score than OOF.</p>\n<p>The high standard deviation does not occur due to the use of AUC but is a consequence of the bias-variance tradeoff. When we use a proper scoring rule, we would find also a high std. For log loss, we have ((-log(p_1) - log(p_2))/2 + (-log(p_3) - log(p_4))/2)/2 = (-log(p_1) - log(p_2) - log(p_3) - log(p_4))/4.</p>\n<p>Of course, it could be a mistake in my notebook or there are some things I didn't pay attention to. If so, I look forward to learning something new.</p>",
      "rawMarkdown": "While your argument makes intuitive sense, I'm not sure I totally agree with that. I've just created a new [notebook](https://www.kaggle.com/lars123/oof-vs-cv) to show some empirical results.\n\nI did a train-test-split on the dataset (according to the public-private LB ratio i.e. 78%, 22%). Then I trained a logistic regression with 5-fold and 60-fold CV and looked at the OOF and CV score. The CV score was in all cases closer to the true score than OOF.\n\nThe high standard deviation does not occur due to the use of AUC but is a consequence of the bias-variance tradeoff. When we use a proper scoring rule, we would find also a high std. For log loss, we have ((-log(p_1) - log(p_2))/2 + (-log(p_3) - log(p_4))/2)/2 = (-log(p_1) - log(p_2) - log(p_3) - log(p_4))/4.\n\nOf course, it could be a mistake in my notebook or there are some things I didn't pay attention to. If so, I look forward to learning something new.",
      "votes": null
    },
    {
      "id": "1529877",
      "postDate": "09/30/2021 18:04:47",
      "content": "<p>I think our dataset is too small to draw any conclusion. I just changed the <code>random_state</code> of <code>train_test_split()</code> in your notebook to 0 and got 5-fold CV AUC (0.588) much closer to test AUC (0.600) than 60-fold CV AUC (0.646). Maybe you should try to test your hypothesis on a larger toy dataset.</p>",
      "rawMarkdown": "I think our dataset is too small to draw any conclusion. I just changed the `random_state` of `train_test_split()` in your notebook to 0 and got 5-fold CV AUC (0.588) much closer to test AUC (0.600) than 60-fold CV AUC (0.646). Maybe you should try to test your hypothesis on a larger toy dataset.",
      "votes": null
    },
    {
      "id": "1529902",
      "postDate": "09/30/2021 18:38:14",
      "content": "<p>Good discussion! I usually prefer the OOF metric over CV average, but I think there is a risk if the metric is rank-based like AUC here.</p>\n<p>When training <code>n</code> different fold models, they can have different output distributions for various reasons (e.g., loss function that doesn't control range, randomness in training, training subset difference, …). Now, when ranking together predictions from different models, slight differences in model-specific prediction distributions could affect the OOF AUC score.</p>\n<p>Another example of a similar rank-mismatch issue could happen when ensembling predictions by averaging. Averaging predictions from two different distributions; one that ranges <code>from 0.2 to 0.6</code> and another that ranges <code>from 0.0 to 1.0</code>, the latter predictions would dominate the ranking metric.</p>",
      "rawMarkdown": "Good discussion! I usually prefer the OOF metric over CV average, but I think there is a risk if the metric is rank-based like AUC here.\n\nWhen training `n` different fold models, they can have different output distributions for various reasons (e.g., loss function that doesn't control range, randomness in training, training subset difference, ...). Now, when ranking together predictions from different models, slight differences in model-specific prediction distributions could affect the OOF AUC score.\n\nAnother example of a similar rank-mismatch issue could happen when ensembling predictions by averaging. Averaging predictions from two different distributions; one that ranges `from 0.2 to 0.6` and another that ranges `from 0.0 to 1.0`, the latter predictions would dominate the ranking metric.",
      "votes": null
    },
    {
      "id": "1529946",
      "postDate": "09/30/2021 19:32:23",
      "content": "<p>You have a point. Let's say we have a distribution X (e.g. public dataset), then we estimate our generalization error on this dataset distribution. But if somebody comes along and gives us a totally different dataset (e.g. private dataset or different train-test split), then our generalization error based on X can be wrong (in this <a href=\"https://www.jmlr.org/papers/volume5/grandvalet04a/grandvalet04a.pdf\" target=\"_blank\">paper</a> they call it \"prediction error\" vs \"expected value of prediction error\"). </p>\n<p><a href=\"https://www.kaggle.com/nanguyen\" target=\"_blank\">@nanguyen</a> you mentioned the problem of different seeds. I added a new section to my public notebook in order to also consider this case. Let X be the dataset distribution. I randomly sample 300 datasets from this distribution X (train_test_split) and then compute the avg AUC (so we have like a \"population AUC\" vs \"sample AUC\"). Still CV with 60 folds is better than 5 folds. Furthermore, the results for OOF are worse.</p>\n<p><a href=\"https://www.kaggle.com/qitvision\" target=\"_blank\">@qitvision</a> I agree that e.g. neural networks / decision trees can cause all kinds of problems because the classifiers are often not calibrated. Decision trees do not like values between 0 and 1. With logistic regression, I have only had good experiences :)</p>",
      "rawMarkdown": "You have a point. Let's say we have a distribution X (e.g. public dataset), then we estimate our generalization error on this dataset distribution. But if somebody comes along and gives us a totally different dataset (e.g. private dataset or different train-test split), then our generalization error based on X can be wrong (in this [paper](https://www.jmlr.org/papers/volume5/grandvalet04a/grandvalet04a.pdf) they call it \"prediction error\" vs \"expected value of prediction error\"). \n\n@nanguyen you mentioned the problem of different seeds. I added a new section to my public notebook in order to also consider this case. Let X be the dataset distribution. I randomly sample 300 datasets from this distribution X (train_test_split) and then compute the avg AUC (so we have like a \"population AUC\" vs \"sample AUC\"). Still CV with 60 folds is better than 5 folds. Furthermore, the results for OOF are worse.\n\n@qitvision I agree that e.g. neural networks / decision trees can cause all kinds of problems because the classifiers are often not calibrated. Decision trees do not like values between 0 and 1. With logistic regression, I have only had good experiences :)",
      "votes": null
    },
    {
      "id": "1530297",
      "postDate": "10/01/2021 04:58:01",
      "content": "<p>This is very interesting. So from your analysis, even if the dataset is small, we can still estimate \"true\" AUC accurately using CV AUC of <code>RepeatedKFold</code> (with <code>n_folds</code> and <code>n_repeats</code> large enough). Am I interpreting your result correctly?</p>",
      "rawMarkdown": "This is very interesting. So from your analysis, even if the dataset is small, we can still estimate \"true\" AUC accurately using CV AUC of `RepeatedKFold` (with `n_folds` and `n_repeats` large enough). Am I interpreting your result correctly?",
      "votes": null
    },
    {
      "id": "1531030",
      "postDate": "10/01/2021 15:07:55",
      "content": "<p>I did not think about using RepeatedKFold, good idea.</p>\n<p>Yes, we get closer to the true population AUC (of a specific dataset subset) with increasing <code>n_folds</code> and <code>n_repeats</code> (but not equal). The authors in the <a href=\"https://lirias.kuleuven.be/bitstream/123456789/346385/3/OnEstimatingModelAccuracy.pdf\" target=\"_blank\">paper</a> did some more tests. The experiment was as follows: draw subsets from 200 samples for training, rest is a fixed test dataset to compute population accuracy (similar to my notebook). Their result is that CV is worse than repeated CV (see page 5, table 2). However, they did not consider the number of folds or AUC.</p>\n<p>I do not want to say that a high <code>n_folds</code> for CV or RepeatedKFold is always necessary because often 5 or 10 folds with regular CV are sufficient. But it does not hurt to use high n_folds+n_repeats when it is computationally feasible (because it can only improve the bias).</p>\n<p>In this <a href=\"https://stats.stackexchange.com/questions/61783/bias-and-variance-in-leave-one-out-vs-k-fold-cross-validation\" target=\"_blank\">post</a>, the authors simulated a dataset with 200 data points.</p>\n<p><img src=\"https://imgur.com/a/lLNcoz2\" alt=\"graph1\"></p>\n<p>A high number of <code>n_folds</code> did not affect negatively the bias.</p>\n<p>Some time ago I repeated the experiment with a real-world dataset (also quite small).</p>\n<p><img src=\"https://imgur.com/a/PLUiVqC\" alt=\"graph2\"></p>\n<p>With increasing <code>n_folds</code> there were fewer fluctuations for the bias. </p>",
      "rawMarkdown": "I did not think about using RepeatedKFold, good idea.\n\nYes, we get closer to the true population AUC (of a specific dataset subset) with increasing `n_folds` and `n_repeats` (but not equal). The authors in the [paper](https://lirias.kuleuven.be/bitstream/123456789/346385/3/OnEstimatingModelAccuracy.pdf) did some more tests. The experiment was as follows: draw subsets from 200 samples for training, rest is a fixed test dataset to compute population accuracy (similar to my notebook). Their result is that CV is worse than repeated CV (see page 5, table 2). However, they did not consider the number of folds or AUC.\n\nI do not want to say that a high `n_folds` for CV or RepeatedKFold is always necessary because often 5 or 10 folds with regular CV are sufficient. But it does not hurt to use high n_folds+n_repeats when it is computationally feasible (because it can only improve the bias).\n\nIn this [post](https://stats.stackexchange.com/questions/61783/bias-and-variance-in-leave-one-out-vs-k-fold-cross-validation), the authors simulated a dataset with 200 data points.\n\n![graph1](https://imgur.com/a/lLNcoz2)\n\nA high number of `n_folds` did not affect negatively the bias.\n\nSome time ago I repeated the experiment with a real-world dataset (also quite small).\n\n![graph2](https://imgur.com/a/PLUiVqC)\n\nWith increasing `n_folds` there were fewer fluctuations for the bias.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1528363,
      "author_name": "nanguyen",
      "author_url": "",
      "post_date": "09/29/2021 15:28:07",
      "content": "<p>I reran your notebook with number of folds changed to 5, got 0.57 validation AUC while public score is still around 0.65. Not perfectly correlated as you suggested but I agree that maybe those metadata features are useful.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1528507,
          "author_name": "qitvision",
          "author_url": "",
          "post_date": "09/29/2021 17:42:05",
          "content": "<p>The std of AUC was very high (0.4+) in the original notebook with 200-fold CV, so I ran it too with 5-fold CV. The average validation AUC was 0.57 and std remained quite high (0.095).</p>\n<p>There could be some correlation between the metadata features and MGMT label but the features don't seem very robust considering the high std of the AUC between folds.</p>\n<p>This is an interesting discovery, and I wonder if CNNs are able to pick up some of the imaging parameter features like ETL and overfit to it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1528685,
          "author_name": "lars123",
          "author_url": "",
          "post_date": "09/29/2021 21:55:27",
          "content": "<p>By decreasing the number of folds (e.g. K=5), we increase the bias of the CV estimator. As K-&gt;N, there is less bias and higher variance. Hence, the big increase in std for K=200. IMHO here a high K is better as it gives a better estimate of the bias (even if std suffers i.e. bias-variance tradeoff). See <a href=\"https://stats.stackexchange.com/questions/61783/bias-and-variance-in-leave-one-out-vs-k-fold-cross-validation\" target=\"_blank\">https://stats.stackexchange.com/questions/61783/bias-and-variance-in-leave-one-out-vs-k-fold-cross-validation</a> Normally, I would also use K=5 or K=10 but the dataset is so small. This is at least as I would interpret the results.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1528588,
      "author_name": "davidbroberts",
      "author_url": "",
      "post_date": "09/29/2021 19:06:30",
      "content": "<p>I don't think theses DICOM values or number of scans(images) have a direct correlation to MGMT for a couple of reasons.  If they do, I suspect it's purely coincidental due to the small sample size.  </p>\n<p>The technologist performing the scans likely did not know the patient's MGMT status (since that usually comes from a post-mortem biopsy). The scan parameters are generally programmed into a protocol by an engineer and not tweaked by the technologist at acquisition time. Protocols differ slightly between machines, companies and radiologists. </p>\n<p>FOV is generally expected to be square for a brain. But if the patient was having a neck and brain scan at the same time, the chosen protocol might use a rectangular FOV. A dedicated brain scan should always use a square FOV.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1528591,
      "author_name": "maxbaugh",
      "author_url": "",
      "post_date": "09/29/2021 19:16:38",
      "content": "<p>I found that one can get to ~59% accuracy with a simple sklearn SVM classifier <em>using the number of images alone</em>, patients with MGMT = 1 tend to have substantially more FLAIR and T2w images.  The number of images was determined by the doctor/radiologist presumably <em>before</em> the methylation status of the tumor was known, therefore the number of slices should be completely decoupled from MGMT status.  This is why it's important to have some understanding of what your data actually is, otherwise spurious correlations can enter.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1528592,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "09/29/2021 19:18:17",
      "content": "<p>The way you are computing your CV AUC is incorrect, you need to do something like this : </p>\n<pre><code>pred_oof = np.zeros(X.shape[0])\nfor fold, (train_index, val_index) in enumerate(StratifiedKFold(n_splits=k).split(X, y)):\n    [...]\n\n    y_pred = regr.predict_proba(X_val)[...,1]    \n    pred_oof[val_index] = y_pred   # save predictions\n\nprint(\"Loss :\", log_loss(y, pred_oof))\nprint(\"AUC :\", roc_auc_score(y, pred_oof))   # compute the AUC on the whole samples\n</code></pre>\n<p>Which results in 0.55 CV and probably means the 0.6+ lb is just noise.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1528676,
          "author_name": "lars123",
          "author_url": "",
          "post_date": "09/29/2021 21:46:57",
          "content": "<p>I think the way you would compute it would result into an OOF score and not a CV score in the <a href=\"https://www.jmlr.org/papers/volume5/grandvalet04a/grandvalet04a.pdf\" target=\"_blank\">traditional ML sense</a>. <a href=\"https://scikit-learn.org/stable/modules/cross_validation.html#cross-validation\" target=\"_blank\">Sklearn</a> also computes it in the same way as far as I can see. Is there any advantage to using regular CV?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1529144,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "09/30/2021 07:22:53",
          "content": "<p>The oof score is more reliable for metrics such as the AUC (especially since you are using 200 folds) - and I believe it is common practice for most Kaggle competitions. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1529265,
          "author_name": "nanguyen",
          "author_url": "",
          "post_date": "09/30/2021 08:58:12",
          "content": "<p>I agree that oof score is common practice for kaggle. But I haven't read anywhere that oof score is more reliable than his method. Actually I think <a href=\"https://www.kaggle.com/lars123\" target=\"_blank\">@lars123</a>'s traditional method is more reliable because it can also calculate the variance.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1529334,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "09/30/2021 10:00:24",
          "content": "<p>The reliability comes from the AUC. For instance using 4 samples and <code>k = 2</code> folds : </p>\n<pre><code>auc_avg = (auc([x1, x2], [y1, y2]) + auc([x3, x4], [y3, y4])) / 2 \nauc_oof = auc([x1, x2, x3, x4], [y1, y2, y3, y4])\n</code></pre>\n<p><br>\nYou don't actually have the following : <code>auc_avg = avg_oof</code></p>\n<p>And in fact (I think - small explaination bellow) you have : <code>auc_avg &gt;= avg_oof</code></p>\n<p>Let's say <code>y1 = 0, y2 = 1, y3 = 0, y4 = 1</code> : <br>\nTo have  <code>auc_oof = 1</code> you need to find <code>x1 &lt;= x3 &lt; x2 &lt;= x4</code><br>\nTo have <code>auc_avg = 1</code> you need to find <code>x1 &lt; x2 and x3 &lt; x4</code> which is much easier.</p>\n<p>In the test set the computation is made on the <code>n</code> samples directly, and is not the average of <code>k</code> <code>n/k</code> AUCs.</p>\n<p>The AUC is a rank based metric, and its stability benefits from being computed on more samples. <br>\nFor instance, <a href=\"https://www.kaggle.com/lars123\" target=\"_blank\">@lars123</a>'s notebook shows a 0.4+ variance for a 0.6+ auc which is huge. <br>\nThis is because he computes 200 aucs on ~3 samples. <br>\nSimply reducing the number of folds as you did reduces the variance a lot, but also the cv score. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1529800,
          "author_name": "lars123",
          "author_url": "",
          "post_date": "09/30/2021 16:57:51",
          "content": "<p>While your argument makes intuitive sense, I'm not sure I totally agree with that. I've just created a new <a href=\"https://www.kaggle.com/lars123/oof-vs-cv\" target=\"_blank\">notebook</a> to show some empirical results.</p>\n<p>I did a train-test-split on the dataset (according to the public-private LB ratio i.e. 78%, 22%). Then I trained a logistic regression with 5-fold and 60-fold CV and looked at the OOF and CV score. The CV score was in all cases closer to the true score than OOF.</p>\n<p>The high standard deviation does not occur due to the use of AUC but is a consequence of the bias-variance tradeoff. When we use a proper scoring rule, we would find also a high std. For log loss, we have ((-log(p_1) - log(p_2))/2 + (-log(p_3) - log(p_4))/2)/2 = (-log(p_1) - log(p_2) - log(p_3) - log(p_4))/4.</p>\n<p>Of course, it could be a mistake in my notebook or there are some things I didn't pay attention to. If so, I look forward to learning something new.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1529877,
          "author_name": "nanguyen",
          "author_url": "",
          "post_date": "09/30/2021 18:04:47",
          "content": "<p>I think our dataset is too small to draw any conclusion. I just changed the <code>random_state</code> of <code>train_test_split()</code> in your notebook to 0 and got 5-fold CV AUC (0.588) much closer to test AUC (0.600) than 60-fold CV AUC (0.646). Maybe you should try to test your hypothesis on a larger toy dataset.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1529902,
          "author_name": "qitvision",
          "author_url": "",
          "post_date": "09/30/2021 18:38:14",
          "content": "<p>Good discussion! I usually prefer the OOF metric over CV average, but I think there is a risk if the metric is rank-based like AUC here.</p>\n<p>When training <code>n</code> different fold models, they can have different output distributions for various reasons (e.g., loss function that doesn't control range, randomness in training, training subset difference, …). Now, when ranking together predictions from different models, slight differences in model-specific prediction distributions could affect the OOF AUC score.</p>\n<p>Another example of a similar rank-mismatch issue could happen when ensembling predictions by averaging. Averaging predictions from two different distributions; one that ranges <code>from 0.2 to 0.6</code> and another that ranges <code>from 0.0 to 1.0</code>, the latter predictions would dominate the ranking metric.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1529946,
          "author_name": "lars123",
          "author_url": "",
          "post_date": "09/30/2021 19:32:23",
          "content": "<p>You have a point. Let's say we have a distribution X (e.g. public dataset), then we estimate our generalization error on this dataset distribution. But if somebody comes along and gives us a totally different dataset (e.g. private dataset or different train-test split), then our generalization error based on X can be wrong (in this <a href=\"https://www.jmlr.org/papers/volume5/grandvalet04a/grandvalet04a.pdf\" target=\"_blank\">paper</a> they call it \"prediction error\" vs \"expected value of prediction error\"). </p>\n<p><a href=\"https://www.kaggle.com/nanguyen\" target=\"_blank\">@nanguyen</a> you mentioned the problem of different seeds. I added a new section to my public notebook in order to also consider this case. Let X be the dataset distribution. I randomly sample 300 datasets from this distribution X (train_test_split) and then compute the avg AUC (so we have like a \"population AUC\" vs \"sample AUC\"). Still CV with 60 folds is better than 5 folds. Furthermore, the results for OOF are worse.</p>\n<p><a href=\"https://www.kaggle.com/qitvision\" target=\"_blank\">@qitvision</a> I agree that e.g. neural networks / decision trees can cause all kinds of problems because the classifiers are often not calibrated. Decision trees do not like values between 0 and 1. With logistic regression, I have only had good experiences :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1530297,
          "author_name": "nanguyen",
          "author_url": "",
          "post_date": "10/01/2021 04:58:01",
          "content": "<p>This is very interesting. So from your analysis, even if the dataset is small, we can still estimate \"true\" AUC accurately using CV AUC of <code>RepeatedKFold</code> (with <code>n_folds</code> and <code>n_repeats</code> large enough). Am I interpreting your result correctly?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1531030,
          "author_name": "lars123",
          "author_url": "",
          "post_date": "10/01/2021 15:07:55",
          "content": "<p>I did not think about using RepeatedKFold, good idea.</p>\n<p>Yes, we get closer to the true population AUC (of a specific dataset subset) with increasing <code>n_folds</code> and <code>n_repeats</code> (but not equal). The authors in the <a href=\"https://lirias.kuleuven.be/bitstream/123456789/346385/3/OnEstimatingModelAccuracy.pdf\" target=\"_blank\">paper</a> did some more tests. The experiment was as follows: draw subsets from 200 samples for training, rest is a fixed test dataset to compute population accuracy (similar to my notebook). Their result is that CV is worse than repeated CV (see page 5, table 2). However, they did not consider the number of folds or AUC.</p>\n<p>I do not want to say that a high <code>n_folds</code> for CV or RepeatedKFold is always necessary because often 5 or 10 folds with regular CV are sufficient. But it does not hurt to use high n_folds+n_repeats when it is computationally feasible (because it can only improve the bias).</p>\n<p>In this <a href=\"https://stats.stackexchange.com/questions/61783/bias-and-variance-in-leave-one-out-vs-k-fold-cross-validation\" target=\"_blank\">post</a>, the authors simulated a dataset with 200 data points.</p>\n<p><img src=\"https://imgur.com/a/lLNcoz2\" alt=\"graph1\"></p>\n<p>A high number of <code>n_folds</code> did not affect negatively the bias.</p>\n<p>Some time ago I repeated the experiment with a real-world dataset (also quite small).</p>\n<p><img src=\"https://imgur.com/a/PLUiVqC\" alt=\"graph2\"></p>\n<p>With increasing <code>n_folds</code> there were fewer fluctuations for the bias. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1528646,
      "author_name": "richardepstein",
      "author_url": "",
      "post_date": "09/29/2021 20:55:45",
      "content": "<p>One possible explanation for the correlation with imaging protocol: </p>\n<p>If the data came from different sites, some sites might preferentially see more MGMT+ patients, and those sites happen to have the protocols we are noticing. </p>\n<p>So we are separating the data based on scanning site and for some reason different sites have a different mix of patients. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1528225": "I found out that the metadata is surprisingly good for training a classifier. Since I highly doubt that the use of this data is allowed, I decided to make this information public.\n\nThe MRI parameters correlate with MGMT promoter methylation. This means that the people who configured the MRI scanner or preprocessed the data leaked some information into the data. I found three parameters that are good indicators of MGMT.\n\n(1) The echo train length (ETL) determines the number of echoes within a repetition time. An increase of ETL introduces more T2 decay in the image and decreases the scan time. \n\n![ETL](https://i.imgur.com/cyAloEV.jpeg)\n\n(2) The phase field of view can also be used to reduce scan time. \n\n![Phase field of view](https://i.imgur.com/LAbzKnJ.jpg)\n\n(3) Number of scans in folders.\n\nI trained a logistic regression and got a CV of 0.6575 and LB of 0.641. LB and CV are perfectly correlated. However, I do not know if private data is also affected.\n\nThe notebook can be found [here](https://www.kaggle.com/lars123/leak-in-metadata).",
    "1528363": "I reran your notebook with number of folds changed to 5, got 0.57 validation AUC while public score is still around 0.65. Not perfectly correlated as you suggested but I agree that maybe those metadata features are useful.",
    "1528507": "The std of AUC was very high (0.4+) in the original notebook with 200-fold CV, so I ran it too with 5-fold CV. The average validation AUC was 0.57 and std remained quite high (0.095).\n\nThere could be some correlation between the metadata features and MGMT label but the features don't seem very robust considering the high std of the AUC between folds.\n\nThis is an interesting discovery, and I wonder if CNNs are able to pick up some of the imaging parameter features like ETL and overfit to it.",
    "1528588": "I don't think theses DICOM values or number of scans(images) have a direct correlation to MGMT for a couple of reasons.  If they do, I suspect it's purely coincidental due to the small sample size.  \n\nThe technologist performing the scans likely did not know the patient's MGMT status (since that usually comes from a post-mortem biopsy). The scan parameters are generally programmed into a protocol by an engineer and not tweaked by the technologist at acquisition time. Protocols differ slightly between machines, companies and radiologists. \n\nFOV is generally expected to be square for a brain. But if the patient was having a neck and brain scan at the same time, the chosen protocol might use a rectangular FOV. A dedicated brain scan should always use a square FOV.",
    "1528591": "I found that one can get to ~59% accuracy with a simple sklearn SVM classifier *using the number of images alone*, patients with MGMT = 1 tend to have substantially more FLAIR and T2w images.  The number of images was determined by the doctor/radiologist presumably *before* the methylation status of the tumor was known, therefore the number of slices should be completely decoupled from MGMT status.  This is why it's important to have some understanding of what your data actually is, otherwise spurious correlations can enter.",
    "1528592": "The way you are computing your CV AUC is incorrect, you need to do something like this : \n\n```\npred_oof = np.zeros(X.shape[0])\nfor fold, (train_index, val_index) in enumerate(StratifiedKFold(n_splits=k).split(X, y)):\n    [...]\n    \n    y_pred = regr.predict_proba(X_val)[...,1]    \n    pred_oof[val_index] = y_pred   # save predictions\n\nprint(\"Loss :\", log_loss(y, pred_oof))\nprint(\"AUC :\", roc_auc_score(y, pred_oof))   # compute the AUC on the whole samples\n```\n\n\nWhich results in 0.55 CV and probably means the 0.6+ lb is just noise.",
    "1528646": "One possible explanation for the correlation with imaging protocol: \n\nIf the data came from different sites, some sites might preferentially see more MGMT+ patients, and those sites happen to have the protocols we are noticing. \n\nSo we are separating the data based on scanning site and for some reason different sites have a different mix of patients.",
    "1528676": "I think the way you would compute it would result into an OOF score and not a CV score in the [traditional ML sense](https://www.jmlr.org/papers/volume5/grandvalet04a/grandvalet04a.pdf). [Sklearn](https://scikit-learn.org/stable/modules/cross_validation.html#cross-validation) also computes it in the same way as far as I can see. Is there any advantage to using regular CV?",
    "1528685": "By decreasing the number of folds (e.g. K=5), we increase the bias of the CV estimator. As K->N, there is less bias and higher variance. Hence, the big increase in std for K=200. IMHO here a high K is better as it gives a better estimate of the bias (even if std suffers i.e. bias-variance tradeoff). See https://stats.stackexchange.com/questions/61783/bias-and-variance-in-leave-one-out-vs-k-fold-cross-validation Normally, I would also use K=5 or K=10 but the dataset is so small. This is at least as I would interpret the results.",
    "1529144": "The oof score is more reliable for metrics such as the AUC (especially since you are using 200 folds) - and I believe it is common practice for most Kaggle competitions.",
    "1529265": "I agree that oof score is common practice for kaggle. But I haven't read anywhere that oof score is more reliable than his method. Actually I think @lars123's traditional method is more reliable because it can also calculate the variance.",
    "1529334": "The reliability comes from the AUC. For instance using 4 samples and `k = 2` folds : \n```\nauc_avg = (auc([x1, x2], [y1, y2]) + auc([x3, x4], [y3, y4])) / 2 \nauc_oof = auc([x1, x2, x3, x4], [y1, y2, y3, y4])\n``` \nYou don't actually have the following : `auc_avg = avg_oof`\n\nAnd in fact (I think - small explaination bellow) you have : `auc_avg >= avg_oof`\n\nLet's say `y1 = 0, y2 = 1, y3 = 0, y4 = 1` : \nTo have  `auc_oof = 1` you need to find `x1 <= x3 < x2 <= x4`\nTo have `auc_avg = 1` you need to find `x1 < x2 and x3 < x4` which is much easier.\n\nIn the test set the computation is made on the `n` samples directly, and is not the average of `k` `n/k` AUCs.\n\n\nThe AUC is a rank based metric, and its stability benefits from being computed on more samples. \nFor instance, @lars123's notebook shows a 0.4+ variance for a 0.6+ auc which is huge. \nThis is because he computes 200 aucs on ~3 samples. \nSimply reducing the number of folds as you did reduces the variance a lot, but also the cv score.",
    "1529800": "While your argument makes intuitive sense, I'm not sure I totally agree with that. I've just created a new [notebook](https://www.kaggle.com/lars123/oof-vs-cv) to show some empirical results.\n\nI did a train-test-split on the dataset (according to the public-private LB ratio i.e. 78%, 22%). Then I trained a logistic regression with 5-fold and 60-fold CV and looked at the OOF and CV score. The CV score was in all cases closer to the true score than OOF.\n\nThe high standard deviation does not occur due to the use of AUC but is a consequence of the bias-variance tradeoff. When we use a proper scoring rule, we would find also a high std. For log loss, we have ((-log(p_1) - log(p_2))/2 + (-log(p_3) - log(p_4))/2)/2 = (-log(p_1) - log(p_2) - log(p_3) - log(p_4))/4.\n\nOf course, it could be a mistake in my notebook or there are some things I didn't pay attention to. If so, I look forward to learning something new.",
    "1529877": "I think our dataset is too small to draw any conclusion. I just changed the `random_state` of `train_test_split()` in your notebook to 0 and got 5-fold CV AUC (0.588) much closer to test AUC (0.600) than 60-fold CV AUC (0.646). Maybe you should try to test your hypothesis on a larger toy dataset.",
    "1529902": "Good discussion! I usually prefer the OOF metric over CV average, but I think there is a risk if the metric is rank-based like AUC here.\n\nWhen training `n` different fold models, they can have different output distributions for various reasons (e.g., loss function that doesn't control range, randomness in training, training subset difference, ...). Now, when ranking together predictions from different models, slight differences in model-specific prediction distributions could affect the OOF AUC score.\n\nAnother example of a similar rank-mismatch issue could happen when ensembling predictions by averaging. Averaging predictions from two different distributions; one that ranges `from 0.2 to 0.6` and another that ranges `from 0.0 to 1.0`, the latter predictions would dominate the ranking metric.",
    "1529946": "You have a point. Let's say we have a distribution X (e.g. public dataset), then we estimate our generalization error on this dataset distribution. But if somebody comes along and gives us a totally different dataset (e.g. private dataset or different train-test split), then our generalization error based on X can be wrong (in this [paper](https://www.jmlr.org/papers/volume5/grandvalet04a/grandvalet04a.pdf) they call it \"prediction error\" vs \"expected value of prediction error\"). \n\n@nanguyen you mentioned the problem of different seeds. I added a new section to my public notebook in order to also consider this case. Let X be the dataset distribution. I randomly sample 300 datasets from this distribution X (train_test_split) and then compute the avg AUC (so we have like a \"population AUC\" vs \"sample AUC\"). Still CV with 60 folds is better than 5 folds. Furthermore, the results for OOF are worse.\n\n@qitvision I agree that e.g. neural networks / decision trees can cause all kinds of problems because the classifiers are often not calibrated. Decision trees do not like values between 0 and 1. With logistic regression, I have only had good experiences :)",
    "1530297": "This is very interesting. So from your analysis, even if the dataset is small, we can still estimate \"true\" AUC accurately using CV AUC of `RepeatedKFold` (with `n_folds` and `n_repeats` large enough). Am I interpreting your result correctly?",
    "1531030": "I did not think about using RepeatedKFold, good idea.\n\nYes, we get closer to the true population AUC (of a specific dataset subset) with increasing `n_folds` and `n_repeats` (but not equal). The authors in the [paper](https://lirias.kuleuven.be/bitstream/123456789/346385/3/OnEstimatingModelAccuracy.pdf) did some more tests. The experiment was as follows: draw subsets from 200 samples for training, rest is a fixed test dataset to compute population accuracy (similar to my notebook). Their result is that CV is worse than repeated CV (see page 5, table 2). However, they did not consider the number of folds or AUC.\n\nI do not want to say that a high `n_folds` for CV or RepeatedKFold is always necessary because often 5 or 10 folds with regular CV are sufficient. But it does not hurt to use high n_folds+n_repeats when it is computationally feasible (because it can only improve the bias).\n\nIn this [post](https://stats.stackexchange.com/questions/61783/bias-and-variance-in-leave-one-out-vs-k-fold-cross-validation), the authors simulated a dataset with 200 data points.\n\n![graph1](https://imgur.com/a/lLNcoz2)\n\nA high number of `n_folds` did not affect negatively the bias.\n\nSome time ago I repeated the experiment with a real-world dataset (also quite small).\n\n![graph2](https://imgur.com/a/PLUiVqC)\n\nWith increasing `n_folds` there were fewer fluctuations for the bias."
  },
  "source": "meta"
}