{
  "id": 290235,
  "title": "A cross-validation strategy: subsequences",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/290235",
  "author_name": "",
  "post_date": "2021-11-23T17:15:43.910357400Z",
  "votes": 37,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hello community,</p>\n<p>You have probably already discovered that there are 3 videos and that that makes cross-validation a bit weird, since this means a 3-fold CV or a weird 66-33 split.</p>\n<p><strong>Sequences</strong>, on the other hand, are not much better: there are only 20 and they are quite dissimilar in their sizes.</p>\n<p>I have come up with a \"subsequence\" concept, which creates 137 \"splitting units\".</p>\n<p>A <strong>subsequence</strong> is a piece of a sequence when objects are continually present.<br>\nOr the opposite, a continuous piece when objects are not present.</p>\n<p>&nbsp;</p>\n<p>Let's see an example. Consider the sequence <code>A</code> with the following frames:</p>\n<ul>\n<li><code>1-20</code> - No annotations present</li>\n<li><code>21-30</code> - Annotations present</li>\n<li><code>31-60</code> - No annotations</li>\n<li><code>61-80</code> - Annotations present</li>\n</ul>\n<p>In this case, we say that the sequence <code>A</code> has <code>4</code> subsequences (<code>1-20</code>, <code>21-30</code>, <code>31-60</code>, <code>61-80</code>).</p>\n<p>&nbsp;</p>\n<p>A subsequence seems to me like the minimal atom for ensuring no leaks happen between train and test.</p>\n<p>I have gone in-depth with this approach and created common splits with it in <strong>this notebook</strong>:</p>\n<h3><a href=\"https://www.kaggle.com/julian3833/reef-cv-strategy-subsequences\" target=\"_blank\">🐠 Reef - CV strategy: subsequences!</a></h3>\n<p>The resulting dataframes are saved in <strong>this dataset</strong>: </p>\n<h3><a href=\"https://www.kaggle.com/julian3833/reef-cv-strategy-subsequences-dataframes/\" target=\"_blank\">reef-cv-strategy-subsequences-dataframes</a></h3>\n<h4>There are train-validation splits with 1%, 5%, 10% and 20% validation sizes and 5 and 10 stratified folds for cross-validation that are ready to use.</h4>\n<p>I'm planning to adapt my notebooks to work with these splits soon. Hopefully you will find them useful.</p>\n<h3>On the other hand: have you thought of other approaches? I could think of this one only, but I imagine others are possible as well and I would love to hear about them.</h3>",
  "messages": [
    {
      "id": "1593146",
      "postDate": "11/23/2021 17:15:43",
      "content": "<p>Hello community,</p>\n<p>You have probably already discovered that there are 3 videos and that that makes cross-validation a bit weird, since this means a 3-fold CV or a weird 66-33 split.</p>\n<p><strong>Sequences</strong>, on the other hand, are not much better: there are only 20 and they are quite dissimilar in their sizes.</p>\n<p>I have come up with a \"subsequence\" concept, which creates 137 \"splitting units\".</p>\n<p>A <strong>subsequence</strong> is a piece of a sequence when objects are continually present.<br>\nOr the opposite, a continuous piece when objects are not present.</p>\n<p>&nbsp;</p>\n<p>Let's see an example. Consider the sequence <code>A</code> with the following frames:</p>\n<ul>\n<li><code>1-20</code> - No annotations present</li>\n<li><code>21-30</code> - Annotations present</li>\n<li><code>31-60</code> - No annotations</li>\n<li><code>61-80</code> - Annotations present</li>\n</ul>\n<p>In this case, we say that the sequence <code>A</code> has <code>4</code> subsequences (<code>1-20</code>, <code>21-30</code>, <code>31-60</code>, <code>61-80</code>).</p>\n<p>&nbsp;</p>\n<p>A subsequence seems to me like the minimal atom for ensuring no leaks happen between train and test.</p>\n<p>I have gone in-depth with this approach and created common splits with it in <strong>this notebook</strong>:</p>\n<h3><a href=\"https://www.kaggle.com/julian3833/reef-cv-strategy-subsequences\" target=\"_blank\">🐠 Reef - CV strategy: subsequences!</a></h3>\n<p>The resulting dataframes are saved in <strong>this dataset</strong>: </p>\n<h3><a href=\"https://www.kaggle.com/julian3833/reef-cv-strategy-subsequences-dataframes/\" target=\"_blank\">reef-cv-strategy-subsequences-dataframes</a></h3>\n<h4>There are train-validation splits with 1%, 5%, 10% and 20% validation sizes and 5 and 10 stratified folds for cross-validation that are ready to use.</h4>\n<p>I'm planning to adapt my notebooks to work with these splits soon. Hopefully you will find them useful.</p>\n<h3>On the other hand: have you thought of other approaches? I could think of this one only, but I imagine others are possible as well and I would love to hear about them.</h3>",
      "rawMarkdown": "Hello community,\n\nYou have probably already discovered that there are 3 videos and that that makes cross-validation a bit weird, since this means a 3-fold CV or a weird 66-33 split.\n\n**Sequences**, on the other hand, are not much better: there are only 20 and they are quite dissimilar in their sizes.\n\nI have come up with a \"subsequence\" concept, which creates 137 \"splitting units\".\n\nA **subsequence** is a piece of a sequence when objects are continually present.\nOr the opposite, a continuous piece when objects are not present.\n\n&nbsp;\n\nLet's see an example. Consider the sequence `A` with the following frames:\n* `1-20` - No annotations present\n* `21-30` - Annotations present\n* `31-60` - No annotations\n* `61-80` - Annotations present\n\nIn this case, we say that the sequence `A` has `4` subsequences (`1-20`, `21-30`, `31-60`, `61-80`).\n\n&nbsp;\n\n\nA subsequence seems to me like the minimal atom for ensuring no leaks happen between train and test.\n\nI have gone in-depth with this approach and created common splits with it in **this notebook**:\n### [🐠 Reef - CV strategy: subsequences!](https://www.kaggle.com/julian3833/reef-cv-strategy-subsequences)\n\nThe resulting dataframes are saved in **this dataset**: \n### [reef-cv-strategy-subsequences-dataframes](https://www.kaggle.com/julian3833/reef-cv-strategy-subsequences-dataframes/)\n\n\n#### There are train-validation splits with 1%, 5%, 10% and 20% validation sizes and 5 and 10 stratified folds for cross-validation that are ready to use.\n\nI'm planning to adapt my notebooks to work with these splits soon. Hopefully you will find them useful.\n\n### On the other hand: have you thought of other approaches? I could think of this one only, but I imagine others are possible as well and I would love to hear about them.",
      "votes": null
    },
    {
      "id": "1593821",
      "postDate": "11/24/2021 09:51:59",
      "content": "<p>Good approach. but what about the edges of that sequence?<br>\nLet say you have a video and divide it into sequences <code>[1-20, 21-40, 41-60, 61-80, 81-100]</code> or <code>[A, B, C, D, E, F]</code>. So, you will feed A, C, E into the training process, and B, D, F for validation, right?</p>\n<p>But what about the similarity between A[19] and B[0]? Subsequence frames will be almost the same. Will overfitted model predict the head-tail of the validation sequence correctly?</p>",
      "rawMarkdown": "Good approach. but what about the edges of that sequence?\nLet say you have a video and divide it into sequences `[1-20, 21-40, 41-60, 61-80, 81-100]` or `[A, B, C, D, E, F]`. So, you will feed A, C, E into the training process, and B, D, F for validation, right?\n\nBut what about the similarity between A[19] and B[0]? Subsequence frames will be almost the same. Will overfitted model predict the head-tail of the validation sequence correctly?",
      "votes": null
    },
    {
      "id": "1594453",
      "postDate": "11/24/2021 21:03:15",
      "content": "<p>Hi Mykola, thanks for commenting.</p>\n<p>The edges are kind of covered by the fact that two consecutive subsequences cannot have both objects. So if A has objects in it, then B doesn't. Only C will have objects, but they will appear in a region far from A. </p>\n<p>Anyway, A[19] and B[0] are continuous frames as you correctly stated, so it is still might be somehow problematic. I think this problem is minor though, but it's just an intuition. Putting continuous subsequences in the same bag (aka training or validation) can strongly reduce these edge cases, but in that case, intra-training and intra-validation variance will be reduced as well (the subsequence ids in the data frames can be used to simply create different splits based on them).</p>",
      "rawMarkdown": "Hi Mykola, thanks for commenting.\n\nThe edges are kind of covered by the fact that two consecutive subsequences cannot have both objects. So if A has objects in it, then B doesn't. Only C will have objects, but they will appear in a region far from A. \n\nAnyway, A[19] and B[0] are continuous frames as you correctly stated, so it is still might be somehow problematic. I think this problem is minor though, but it's just an intuition. Putting continuous subsequences in the same bag (aka training or validation) can strongly reduce these edge cases, but in that case, intra-training and intra-validation variance will be reduced as well (the subsequence ids in the data frames can be used to simply create different splits based on them).",
      "votes": null
    },
    {
      "id": "1660076",
      "postDate": "01/22/2022 11:30:38",
      "content": "<p>I just come up with that sequence type cross validation, and discover that you have already published it! Thanks bro.</p>",
      "rawMarkdown": "I just come up with that sequence type cross validation, and discover that you have already published it! Thanks bro.",
      "votes": null
    },
    {
      "id": "1669904",
      "postDate": "01/31/2022 03:54:46",
      "content": "<p>Great work. Thanks for sharing, this is helpful. I'm using your sub sequences. </p>\n<p>Have you considered <code>sklearn.model_selection.StratifiedGroupKFold</code>? </p>\n<pre><code>from sklearn.model_selection import StratifiedGroupKFold\nsgkf = StratifiedGroupKFold(n_splits=5)\nfor fold, (t_idx, v_idx) in enumerate( sgkf.split(df, df.has_annotations, df.subsequence_id) ):\n    df.loc[v_idx,'fold'] = fold\n</code></pre>\n<p>This produces 5 folds where each fold has 4700 frames and each fold has 984 annotated frames.</p>",
      "rawMarkdown": "Great work. Thanks for sharing, this is helpful. I'm using your sub sequences. \n\nHave you considered `sklearn.model_selection.StratifiedGroupKFold`? \n\n    from sklearn.model_selection import StratifiedGroupKFold\n    sgkf = StratifiedGroupKFold(n_splits=5)\n    for fold, (t_idx, v_idx) in enumerate( sgkf.split(df, df.has_annotations, df.subsequence_id) ):\n        df.loc[v_idx,'fold'] = fold\n\nThis produces 5 folds where each fold has 4700 frames and each fold has 984 annotated frames.",
      "votes": null
    },
    {
      "id": "1671971",
      "postDate": "02/01/2022 20:11:52",
      "content": "<p>Nice code! For those who want to use this code on Kaggle - you'll have to update scikit-learn package. You can do it with this code:</p>\n<pre><code>!pip uninstall -q -y scikit-learn\n!pip install -q scikit-learn\n</code></pre>",
      "rawMarkdown": "Nice code! For those who want to use this code on Kaggle - you'll have to update scikit-learn package. You can do it with this code:\n```\n!pip uninstall -q -y scikit-learn\n!pip install -q scikit-learn\n```",
      "votes": null
    },
    {
      "id": "1672579",
      "postDate": "02/02/2022 07:09:16",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> what is the grouping of cv strategy? by sequence?</p>",
      "rawMarkdown": "cdeotte what is the grouping of cv strategy? by sequence?",
      "votes": null
    },
    {
      "id": "1673103",
      "postDate": "02/02/2022 14:12:36",
      "content": "<p>No, group by <code>sub-sequence</code> which are made in notebook that is linked above.</p>",
      "rawMarkdown": "No, group by `sub-sequence` which are made in notebook that is linked above.",
      "votes": null
    },
    {
      "id": "1673187",
      "postDate": "02/02/2022 15:06:07",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , thanks, I overlooked \"df\" :D</p>",
      "rawMarkdown": "cdeotte , thanks, I overlooked \"df\" :D",
      "votes": null
    },
    {
      "id": "1673938",
      "postDate": "02/03/2022 05:29:19",
      "content": "<p>Great work. Thanks for sharing, this is helpful. </p>",
      "rawMarkdown": "Great work. Thanks for sharing, this is helpful.",
      "votes": null
    },
    {
      "id": "1676122",
      "postDate": "02/04/2022 17:34:35",
      "content": "<p>Why the F2 reporting hits high 0.7 to 0.8 range with this subsequence strategy ? I'm using the 0.1 train validation split </p>",
      "rawMarkdown": "Why the F2 reporting hits high 0.7 to 0.8 range with this subsequence strategy ? I'm using the 0.1 train validation split",
      "votes": null
    },
    {
      "id": "1676130",
      "postDate": "02/04/2022 17:44:13",
      "content": "<p>I think it is because when Yolo reports F2, it is computing F2 per image (then average). Whereas Kaggle competition computes F2 per everything at once.</p>\n<p>(I have also have a difference between what Yolo reports and what we get when computing competition metric ourselves).</p>",
      "rawMarkdown": "I think it is because when Yolo reports F2, it is computing F2 per image (then average). Whereas Kaggle competition computes F2 per everything at once.\n\n(I have also have a difference between what Yolo reports and what we get when computing competition metric ourselves).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1593821,
      "author_name": "meowmeowmeowmeowmeow",
      "author_url": "",
      "post_date": "11/24/2021 09:51:59",
      "content": "<p>Good approach. but what about the edges of that sequence?<br>\nLet say you have a video and divide it into sequences <code>[1-20, 21-40, 41-60, 61-80, 81-100]</code> or <code>[A, B, C, D, E, F]</code>. So, you will feed A, C, E into the training process, and B, D, F for validation, right?</p>\n<p>But what about the similarity between A[19] and B[0]? Subsequence frames will be almost the same. Will overfitted model predict the head-tail of the validation sequence correctly?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1594453,
          "author_name": "julian3833",
          "author_url": "",
          "post_date": "11/24/2021 21:03:15",
          "content": "<p>Hi Mykola, thanks for commenting.</p>\n<p>The edges are kind of covered by the fact that two consecutive subsequences cannot have both objects. So if A has objects in it, then B doesn't. Only C will have objects, but they will appear in a region far from A. </p>\n<p>Anyway, A[19] and B[0] are continuous frames as you correctly stated, so it is still might be somehow problematic. I think this problem is minor though, but it's just an intuition. Putting continuous subsequences in the same bag (aka training or validation) can strongly reduce these edge cases, but in that case, intra-training and intra-validation variance will be reduced as well (the subsequence ids in the data frames can be used to simply create different splits based on them).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1660076,
      "author_name": "calvchen",
      "author_url": "",
      "post_date": "01/22/2022 11:30:38",
      "content": "<p>I just come up with that sequence type cross validation, and discover that you have already published it! Thanks bro.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1669904,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "01/31/2022 03:54:46",
      "content": "<p>Great work. Thanks for sharing, this is helpful. I'm using your sub sequences. </p>\n<p>Have you considered <code>sklearn.model_selection.StratifiedGroupKFold</code>? </p>\n<pre><code>from sklearn.model_selection import StratifiedGroupKFold\nsgkf = StratifiedGroupKFold(n_splits=5)\nfor fold, (t_idx, v_idx) in enumerate( sgkf.split(df, df.has_annotations, df.subsequence_id) ):\n    df.loc[v_idx,'fold'] = fold\n</code></pre>\n<p>This produces 5 folds where each fold has 4700 frames and each fold has 984 annotated frames.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1671971,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "02/01/2022 20:11:52",
          "content": "<p>Nice code! For those who want to use this code on Kaggle - you'll have to update scikit-learn package. You can do it with this code:</p>\n<pre><code>!pip uninstall -q -y scikit-learn\n!pip install -q scikit-learn\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1672579,
          "author_name": "projdev",
          "author_url": "",
          "post_date": "02/02/2022 07:09:16",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> what is the grouping of cv strategy? by sequence?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1673103,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/02/2022 14:12:36",
          "content": "<p>No, group by <code>sub-sequence</code> which are made in notebook that is linked above.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1673187,
          "author_name": "projdev",
          "author_url": "",
          "post_date": "02/02/2022 15:06:07",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , thanks, I overlooked \"df\" :D</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1676122,
          "author_name": "nyleve",
          "author_url": "",
          "post_date": "02/04/2022 17:34:35",
          "content": "<p>Why the F2 reporting hits high 0.7 to 0.8 range with this subsequence strategy ? I'm using the 0.1 train validation split </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1676130,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/04/2022 17:44:13",
          "content": "<p>I think it is because when Yolo reports F2, it is computing F2 per image (then average). Whereas Kaggle competition computes F2 per everything at once.</p>\n<p>(I have also have a difference between what Yolo reports and what we get when computing competition metric ourselves).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1673938,
      "author_name": "spectreprediction",
      "author_url": "",
      "post_date": "02/03/2022 05:29:19",
      "content": "<p>Great work. Thanks for sharing, this is helpful. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1593146": "Hello community,\n\nYou have probably already discovered that there are 3 videos and that that makes cross-validation a bit weird, since this means a 3-fold CV or a weird 66-33 split.\n\n**Sequences**, on the other hand, are not much better: there are only 20 and they are quite dissimilar in their sizes.\n\nI have come up with a \"subsequence\" concept, which creates 137 \"splitting units\".\n\nA **subsequence** is a piece of a sequence when objects are continually present.\nOr the opposite, a continuous piece when objects are not present.\n\n&nbsp;\n\nLet's see an example. Consider the sequence `A` with the following frames:\n* `1-20` - No annotations present\n* `21-30` - Annotations present\n* `31-60` - No annotations\n* `61-80` - Annotations present\n\nIn this case, we say that the sequence `A` has `4` subsequences (`1-20`, `21-30`, `31-60`, `61-80`).\n\n&nbsp;\n\n\nA subsequence seems to me like the minimal atom for ensuring no leaks happen between train and test.\n\nI have gone in-depth with this approach and created common splits with it in **this notebook**:\n### [🐠 Reef - CV strategy: subsequences!](https://www.kaggle.com/julian3833/reef-cv-strategy-subsequences)\n\nThe resulting dataframes are saved in **this dataset**: \n### [reef-cv-strategy-subsequences-dataframes](https://www.kaggle.com/julian3833/reef-cv-strategy-subsequences-dataframes/)\n\n\n#### There are train-validation splits with 1%, 5%, 10% and 20% validation sizes and 5 and 10 stratified folds for cross-validation that are ready to use.\n\nI'm planning to adapt my notebooks to work with these splits soon. Hopefully you will find them useful.\n\n### On the other hand: have you thought of other approaches? I could think of this one only, but I imagine others are possible as well and I would love to hear about them.",
    "1593821": "Good approach. but what about the edges of that sequence?\nLet say you have a video and divide it into sequences `[1-20, 21-40, 41-60, 61-80, 81-100]` or `[A, B, C, D, E, F]`. So, you will feed A, C, E into the training process, and B, D, F for validation, right?\n\nBut what about the similarity between A[19] and B[0]? Subsequence frames will be almost the same. Will overfitted model predict the head-tail of the validation sequence correctly?",
    "1594453": "Hi Mykola, thanks for commenting.\n\nThe edges are kind of covered by the fact that two consecutive subsequences cannot have both objects. So if A has objects in it, then B doesn't. Only C will have objects, but they will appear in a region far from A. \n\nAnyway, A[19] and B[0] are continuous frames as you correctly stated, so it is still might be somehow problematic. I think this problem is minor though, but it's just an intuition. Putting continuous subsequences in the same bag (aka training or validation) can strongly reduce these edge cases, but in that case, intra-training and intra-validation variance will be reduced as well (the subsequence ids in the data frames can be used to simply create different splits based on them).",
    "1660076": "I just come up with that sequence type cross validation, and discover that you have already published it! Thanks bro.",
    "1669904": "Great work. Thanks for sharing, this is helpful. I'm using your sub sequences. \n\nHave you considered `sklearn.model_selection.StratifiedGroupKFold`? \n\n    from sklearn.model_selection import StratifiedGroupKFold\n    sgkf = StratifiedGroupKFold(n_splits=5)\n    for fold, (t_idx, v_idx) in enumerate( sgkf.split(df, df.has_annotations, df.subsequence_id) ):\n        df.loc[v_idx,'fold'] = fold\n\nThis produces 5 folds where each fold has 4700 frames and each fold has 984 annotated frames.",
    "1671971": "Nice code! For those who want to use this code on Kaggle - you'll have to update scikit-learn package. You can do it with this code:\n```\n!pip uninstall -q -y scikit-learn\n!pip install -q scikit-learn\n```",
    "1672579": "cdeotte what is the grouping of cv strategy? by sequence?",
    "1673103": "No, group by `sub-sequence` which are made in notebook that is linked above.",
    "1673187": "cdeotte , thanks, I overlooked \"df\" :D",
    "1673938": "Great work. Thanks for sharing, this is helpful.",
    "1676122": "Why the F2 reporting hits high 0.7 to 0.8 range with this subsequence strategy ? I'm using the 0.1 train validation split",
    "1676130": "I think it is because when Yolo reports F2, it is computing F2 per image (then average). Whereas Kaggle competition computes F2 per everything at once.\n\n(I have also have a difference between what Yolo reports and what we get when computing competition metric ourselves)."
  },
  "source": "meta"
}