{
  "id": 67819,
  "title": "Multilabel Stratification Python Package",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/67819",
  "author_name": "",
  "post_date": "2018-10-05T19:09:55.759266300Z",
  "votes": 103,
  "comment_count": 26,
  "views": 0,
  "content": "<p>Participants may find it helpful to balance distributions of multilabel data across splits for cross validation (i.e., stratify the data). Earlier this year I created a Python package called iterative-stratification that aims to accomplish this task for multilabel data: <a href=\"https://github.com/trent-b/iterative-stratification\">https://github.com/trent-b/iterative-stratification</a>. I hope the package may find some utility in this competition.</p>",
  "messages": [
    {
      "id": "399431",
      "postDate": "10/05/2018 19:09:55",
      "content": "<p>Participants may find it helpful to balance distributions of multilabel data across splits for cross validation (i.e., stratify the data). Earlier this year I created a Python package called iterative-stratification that aims to accomplish this task for multilabel data: <a href=\"https://github.com/trent-b/iterative-stratification\">https://github.com/trent-b/iterative-stratification</a>. I hope the package may find some utility in this competition.</p>",
      "rawMarkdown": "Participants may find it helpful to balance distributions of multilabel data across splits for cross validation (i.e., stratify the data). Earlier this year I created a Python package called iterative-stratification that aims to accomplish this task for multilabel data: https://github.com/trent-b/iterative-stratification. I hope the package may find some utility in this competition.",
      "votes": null
    },
    {
      "id": "399655",
      "postDate": "10/06/2018 11:57:23",
      "content": "<p>This is really useful for multi label classification. Thank you very much! </p>",
      "rawMarkdown": "This is really useful for multi label classification. Thank you very much!",
      "votes": null
    },
    {
      "id": "399670",
      "postDate": "10/06/2018 13:02:21",
      "content": "<p>You're welcome. Please don't hesitate to let me know of any questions you have or issues you encounter via my GitHub project page. Thanks!</p>",
      "rawMarkdown": "You're welcome. Please don't hesitate to let me know of any questions you have or issues you encounter via my GitHub project page. Thanks!",
      "votes": null
    },
    {
      "id": "400121",
      "postDate": "10/07/2018 16:33:59",
      "content": "<p><a href=\"https://www.kaggle.com/kmader/rgb-transfer-learning-with-vgg16-for-protein-atlas\">Here</a> is another way to split data:</p>\n\n<pre><code>raw_train_df, valid_df = train_test_split(image_df, \n                 test_size = 0.3, \n                  # hack to make stratification work                  \n                 stratify = image_df['Target'].map(lambda x: x[:3] if '27' not in x else '0'))\n</code></pre>\n\n<p>Does your approch have advantages over this ? </p>",
      "rawMarkdown": "[Here][1] is another way to split data:\n\n    raw_train_df, valid_df = train_test_split(image_df, \n                     test_size = 0.3, \n                      # hack to make stratification work                  \n                     stratify = image_df['Target'].map(lambda x: x[:3] if '27' not in x else '0'))\n\nDoes your approch have advantages over this ? \n\n  [1]: https://www.kaggle.com/kmader/rgb-transfer-learning-with-vgg16-for-protein-atlas",
      "votes": null
    },
    {
      "id": "400153",
      "postDate": "10/07/2018 17:50:53",
      "content": "<p>Good question. It appears that the code you provide performs labelsets-based stratification. This is certainly a valid approach. In <a href=\"http://lpis.csd.auth.gr/publications/sechidis-ecmlpkdd-2011.pdf\">Sechidis et al. (2011)</a>, the authors compare random sampling, labelset-based stratification, and iterative stratification (the method I have implemented) for multiple datasets. There are cases where labelsets-stratification is better than iterative stratification. These cases are typically when the ratio of unique labelsets to the number of samples is small. In cases where the ratio of unique labelsets to the number of samples is not small, iterative stratification is often the better choice.</p>",
      "rawMarkdown": "Good question. It appears that the code you provide performs labelsets-based stratification. This is certainly a valid approach. In [Sechidis et al. (2011)][1], the authors compare random sampling, labelset-based stratification, and iterative stratification (the method I have implemented) for multiple datasets. There are cases where labelsets-stratification is better than iterative stratification. These cases are typically when the ratio of unique labelsets to the number of samples is small. In cases where the ratio of unique labelsets to the number of samples is not small, iterative stratification is often the better choice.\n\n\n  [1]: http://lpis.csd.auth.gr/publications/sechidis-ecmlpkdd-2011.pdf",
      "votes": null
    },
    {
      "id": "400154",
      "postDate": "10/07/2018 17:53:12",
      "content": "<p>Thank you for answer, will try both of them ;)</p>",
      "rawMarkdown": "Thank you for answer, will try both of them ;)",
      "votes": null
    },
    {
      "id": "400549",
      "postDate": "10/08/2018 14:20:48",
      "content": "<p>The approach there (I used it in my notebook) is very much an engineer hack based on gut feelings, not statistical know-how. @Trentb's package looks much better and it would be cool to get this into the standard kaggle kernel toolkit (it would also be useful for the NIH Chest X-Ray data)</p>",
      "rawMarkdown": "The approach there (I used it in my notebook) is very much an engineer hack based on gut feelings, not statistical know-how. @Trentb's package looks much better and it would be cool to get this into the standard kaggle kernel toolkit (it would also be useful for the NIH Chest X-Ray data)",
      "votes": null
    },
    {
      "id": "400822",
      "postDate": "10/09/2018 00:28:14",
      "content": "<p>Thank you for the recommendation. I have submitted a pull request to add iterative-stratification to Kaggle Kernels. In the meantime, adding iterative-stratification as a custom package in the settings of a kernel works for me.</p>",
      "rawMarkdown": "Thank you for the recommendation. I have submitted a pull request to add iterative-stratification to Kaggle Kernels. In the meantime, adding iterative-stratification as a custom package in the settings of a kernel works for me.",
      "votes": null
    },
    {
      "id": "417662",
      "postDate": "11/08/2018 15:46:57",
      "content": "<p>I do not understand why this 'hack' would work:\n<code>stratify = image_df['Target'].map(lambda x: x[:3] if '27' not in x else '0'))</code>\nIt looks like you just take 3 first characters of each label, I don't know why it would make sense. Could someone please elaborate?</p>",
      "rawMarkdown": "I do not understand why this 'hack' would work:\n`stratify = image_df['Target'].map(lambda x: x[:3] if '27' not in x else '0'))`\nIt looks like you just take 3 first characters of each label, I don't know why it would make sense. Could someone please elaborate?",
      "votes": null
    },
    {
      "id": "417682",
      "postDate": "11/08/2018 16:28:56",
      "content": "<p>This looks pretty useful. One questions I have: can it do uneven splits? I don't have the compute power to be running multiple folds, I'd like to do something like train_test_split does where you can specify a test size. I quickly looked through the github and didn't see any options for that.</p>",
      "rawMarkdown": "This looks pretty useful. One questions I have: can it do uneven splits? I don't have the compute power to be running multiple folds, I'd like to do something like train_test_split does where you can specify a test size. I quickly looked through the github and didn't see any options for that.",
      "votes": null
    },
    {
      "id": "417864",
      "postDate": "11/08/2018 22:15:09",
      "content": "<p>if you want for instance 20% test data you can create a 5fold split but only use the first fold. It should ave the good proportion of test / train. If you want 12.5% a 8fold etc...</p>",
      "rawMarkdown": "if you want for instance 20% test data you can create a 5fold split but only use the first fold. It should ave the good proportion of test / train. If you want 12.5% a 8fold etc...",
      "votes": null
    },
    {
      "id": "418908",
      "postDate": "11/10/2018 21:06:35",
      "content": "<p><a href=\"/ldm314\">@ldm314</a> Yes, you can perform an uneven split. You can call <code>MultilabelStratifiedShuffleSplit(n_splits=1, test_size=0.2, random_state=0)</code> if you want 20% of your data to be for testing and 80% for training.</p>",
      "rawMarkdown": "ldm314 Yes, you can perform an uneven split. You can call `MultilabelStratifiedShuffleSplit(n_splits=1, test_size=0.2, random_state=0)` if you want 20% of your data to be for testing and 80% for training.",
      "votes": null
    },
    {
      "id": "418932",
      "postDate": "11/10/2018 22:45:09",
      "content": "<p>Worked great, using it to split the input dataframe. Will see how this splitting method works out, thanks!</p>\n\n<pre>from iterstrat.ml_stratifiers import MultilabelStratifiedShuffleSplit\nimport numpy as np\n\nmsss = MultilabelStratifiedShuffleSplit(n_splits=1, test_size=0.2, random_state=42)\ntrain_df_orig = train_df.copy()\nX = train_df_orig['Id'].tolist()\ny = train_df_orig['target_vec_float'].tolist()\n\nfor train_index, test_index in msss.split(X,y): #it should only do one iteration\n    print(\"TRAIN:\", train_index, \"TEST:\", test_index)\n    train_df = train_df_orig.loc[train_df_orig.index.intersection(train_index)].copy()\n    valid_df = train_df_orig.loc[train_df_orig.index.intersection(test_index)].copy()\ngc.collect()\n</pre>",
      "rawMarkdown": "Worked great, using it to split the input dataframe. Will see how this splitting method works out, thanks!\n\n<pre>from iterstrat.ml_stratifiers import MultilabelStratifiedShuffleSplit\nimport numpy as np\n\nmsss = MultilabelStratifiedShuffleSplit(n_splits=1, test_size=0.2, random_state=42)\ntrain_df_orig = train_df.copy()\nX = train_df_orig['Id'].tolist()\ny = train_df_orig['target_vec_float'].tolist()\n\nfor train_index, test_index in msss.split(X,y): #it should only do one iteration\n    print(\"TRAIN:\", train_index, \"TEST:\", test_index)\n    train_df = train_df_orig.loc[train_df_orig.index.intersection(train_index)].copy()\n    valid_df = train_df_orig.loc[train_df_orig.index.intersection(test_index)].copy()\ngc.collect()\n</pre>",
      "votes": null
    },
    {
      "id": "424083",
      "postDate": "11/19/2018 14:42:09",
      "content": "<p>I got en error: ValueError: Supported target type is: multilabel-indicator. Got 'binary' instead.  Anyone can help?</p>",
      "rawMarkdown": "I got en error: ValueError: Supported target type is: multilabel-indicator. Got 'binary' instead.  Anyone can help?",
      "votes": null
    },
    {
      "id": "424156",
      "postDate": "11/19/2018 16:42:12",
      "content": "<p>Ensure that each target instance is represented as a list/array of the same size and that the lists/arrays consist of 1s and 0s where 1s represent the presence of a subcellular object.</p>",
      "rawMarkdown": "Ensure that each target instance is represented as a list/array of the same size and that the lists/arrays consist of 1s and 0s where 1s represent the presence of a subcellular object.",
      "votes": null
    },
    {
      "id": "424318",
      "postDate": "11/19/2018 22:07:49",
      "content": "<p>Don't want to hijack this thread... but can you tell me the difference between this package and <a href=\"http://scikit.ml/\">scikit-multilearn</a> in terms of features?</p>\n\n<p>Thank you</p>",
      "rawMarkdown": "Don't want to hijack this thread... but can you tell me the difference between this package and [scikit-multilearn](http://scikit.ml/) in terms of features?\n\nThank you",
      "votes": null
    },
    {
      "id": "424332",
      "postDate": "11/19/2018 23:12:17",
      "content": "<p>Thundo, I'm glad you brought scikit-multilearn to my attention. I was not aware of another scikit-learn-compatible package that performs multi-label stratification. Looking at the dates in GitHub, I believe multi-label stratification was added to scikit-multilearn around the same time that I discovered the paper on iterative stratification; hence, this is probably why my Google searches didn't find an existing implementation of iterative stratification. Scikit-multilearn offers more than multi-label stratification; whereas, my package solely implements multi-label stratification. I'm interested to hear any comparisons on the two implementations of iterative stratification.</p>",
      "rawMarkdown": "Thundo, I'm glad you brought scikit-multilearn to my attention. I was not aware of another scikit-learn-compatible package that performs multi-label stratification. Looking at the dates in GitHub, I believe multi-label stratification was added to scikit-multilearn around the same time that I discovered the paper on iterative stratification; hence, this is probably why my Google searches didn't find an existing implementation of iterative stratification. Scikit-multilearn offers more than multi-label stratification; whereas, my package solely implements multi-label stratification. I'm interested to hear any comparisons on the two implementations of iterative stratification.",
      "votes": null
    },
    {
      "id": "424381",
      "postDate": "11/20/2018 02:36:44",
      "content": "<p>Solved. Thank you very much!</p>",
      "rawMarkdown": "Solved. Thank you very much!",
      "votes": null
    },
    {
      "id": "434742",
      "postDate": "12/06/2018 21:53:55",
      "content": "<p>Hi, I have tried both to see what works best. Tested (1)sklearn.model_selection.train_test_split (this did not involve any stratification), (2)skmultilearn.model_selection.iterative_train_test_split, and (3)your implementation. I've also tried to use stratify kwarg for the train_test_split, but found out that it does not work with multilabel examples. I did simple test/train class count ratios on test_size=0.2, then calculated std between the class count ratios of each implementation.\nThis is what I've got:\n(1) std = 0.04475211027333776 (on different rnd state std = 0.04861630879603909\n(2) std = 0.029059979822161946 (on different run I've got 0.06246754109552749 , not sure why, guess it's different rnd state, but there is no option to set the rnd state)\n(3) std = 0.007712835765265891 (changing rnd state produces the same outcome)\nSeems like your technique works best, if a low std is what we want, which I think is the case. \nThe (2) looks pretty broken, but maybe it is designed to handle some stuff which I am not aware of, or it does something desirable what I did not spot. Was looking at the class count distributions, but nothing obvious on why would (2) split it in this uneven way.\nDid not have the balls to check the difference between the (2) and (3) code directly, would definitely clarify stuff, but I think it would give me a proper headache.\nIf I will have some spare time I could try to train on all of those splits to see what gives the best training results.\nFWIW I am rolling with your code, thank you.</p>",
      "rawMarkdown": "Hi, I have tried both to see what works best. Tested (1)sklearn.model_selection.train_test_split (this did not involve any stratification), (2)skmultilearn.model_selection.iterative_train_test_split, and (3)your implementation. I've also tried to use stratify kwarg for the train_test_split, but found out that it does not work with multilabel examples. I did simple test/train class count ratios on test_size=0.2, then calculated std between the class count ratios of each implementation.\nThis is what I've got:\n(1) std = 0.04475211027333776 (on different rnd state std = 0.04861630879603909\n(2) std = 0.029059979822161946 (on different run I've got 0.06246754109552749 , not sure why, guess it's different rnd state, but there is no option to set the rnd state)\n(3) std = 0.007712835765265891 (changing rnd state produces the same outcome)\nSeems like your technique works best, if a low std is what we want, which I think is the case. \nThe (2) looks pretty broken, but maybe it is designed to handle some stuff which I am not aware of, or it does something desirable what I did not spot. Was looking at the class count distributions, but nothing obvious on why would (2) split it in this uneven way.\nDid not have the balls to check the difference between the (2) and (3) code directly, would definitely clarify stuff, but I think it would give me a proper headache.\nIf I will have some spare time I could try to train on all of those splits to see what gives the best training results.\nFWIW I am rolling with your code, thank you.",
      "votes": null
    },
    {
      "id": "434751",
      "postDate": "12/06/2018 22:12:06",
      "content": "<p>About (2) random state... I usually don't set an explicit <code>random_state</code> in the stratificator. However, since sklearn seeds its random state via numpy you can seed everything with <code>np.random.seed(SEED)</code></p>",
      "rawMarkdown": "About (2) random state... I usually don't set an explicit `random_state` in the stratificator. However, since sklearn seeds its random state via numpy you can seed everything with `np.random.seed(SEED)`",
      "votes": null
    },
    {
      "id": "434812",
      "postDate": "12/07/2018 01:34:28",
      "content": "<p>Thank you very much for your suggestion! I don't have too much time to commit to this competition, but I also would like to share my idea. \nMy data loader just iterate through each class indices and feed them evenly into my model(e.g. 2 per class, makes batch size 56), when one class indices run out, just reshuffle it. This requires more data augmentation than usual imo, but not sure.</p>",
      "rawMarkdown": "Thank you very much for your suggestion! I don't have too much time to commit to this competition, but I also would like to share my idea. \nMy data loader just iterate through each class indices and feed them evenly into my model(e.g. 2 per class, makes batch size 56), when one class indices run out, just reshuffle it. This requires more data augmentation than usual imo, but not sure.",
      "votes": null
    },
    {
      "id": "434821",
      "postDate": "12/07/2018 02:14:51",
      "content": "<p>Karl, thank you for running a comparison test! I'm excited that my code produced such a relatively small std. I'll look into why the rnd state didn't seem to have an effect. In the meantime, something you can try is to set <code>n_splits</code> to 2 (somewhat of a misnomer for scikit-learn's StratifedShuffleSplit) to produce two sets of train/test splits at 80%/20%. They will likely be different as if you had run it twice with <code>n_splits</code> set to 1 and different rnd states.</p>",
      "rawMarkdown": "Karl, thank you for running a comparison test! I'm excited that my code produced such a relatively small std. I'll look into why the rnd state didn't seem to have an effect. In the meantime, something you can try is to set ```n_splits``` to 2 (somewhat of a misnomer for scikit-learn's StratifedShuffleSplit) to produce two sets of train/test splits at 80%/20%. They will likely be different as if you had run it twice with ```n_splits``` set to 1 and different rnd states.",
      "votes": null
    },
    {
      "id": "434908",
      "postDate": "12/07/2018 06:13:43",
      "content": "<p>thanks a lot, very helpful. im using MultilabelStratifiedShuffleSplit now.</p>",
      "rawMarkdown": "thanks a lot, very helpful. im using MultilabelStratifiedShuffleSplit now.",
      "votes": null
    },
    {
      "id": "435531",
      "postDate": "12/08/2018 07:14:49",
      "content": "<p>Hi Karl, <br>\nthanks for running this interesting test! <br>\nDid you or anyone else try different split approaches and reported result in terms of score? (average score of folds). <br>\nThanks in advance</p>",
      "rawMarkdown": "Hi Karl,   \nthanks for running this interesting test!  \nDid you or anyone else try different split approaches and reported result in terms of score? (average score of folds).   \nThanks in advance",
      "votes": null
    },
    {
      "id": "438702",
      "postDate": "12/14/2018 03:34:10",
      "content": "<p>the map function just split all the multilabel into 74 classes. If don't do this, the classes will be very very large</p>",
      "rawMarkdown": "the map function just split all the multilabel into 74 classes. If don't do this, the classes will be very very large",
      "votes": null
    },
    {
      "id": "458549",
      "postDate": "01/19/2019 22:37:13",
      "content": "<p>It seems Karl used <code>skmultilearn.model_selection.iterative_train_test_split</code>  which cannot split into multiple CV splits. it just splits into one train/test set.) <strong>Update:</strong> It seems now this method is <code>skmultilearn.model_selection.iterative_stratification.iterative_train_test_split(X, y, test_size)</code>  See the bottom of the page:\n<a href=\"http://scikit.ml/api/skmultilearn.model_selection.iterative_stratification.html#skmultilearn.model_selection.iterative_stratification.iterative_train_test_split\">http://scikit.ml/api/skmultilearn.model_selection.iterative_stratification.html#skmultilearn.model_selection.iterative_stratification.iterative_train_test_split</a></p>\n\n<p>I checked that they have CV split method too:\n<a href=\"http://scikit.ml/api/skmultilearn.model_selection.iterative_stratification.html\">http://scikit.ml/api/skmultilearn.model_selection.iterative_stratification.html</a></p>\n\n<p>Both methods mentioned in this index page:\n<a href=\"http://scikit.ml/genindex.html\">http://scikit.ml/genindex.html</a></p>\n\n<p>I made more extensive std tests with:\n1. Your implementation\nmskf = MultilabelStratifiedKFold(n_splits=5, shuffle=True, random_state=0)</p>\n\n<ol>\n<li>Scikit implementation\nmskf = IterativeStratification(n_splits=5, order=1)</li>\n</ol>\n\n<p><strong>Tests:</strong>\n*<em>1. Std of test/train splits</em>*</p>\n\n<p>Here is a useful comparison example can be made on the ratios of test label count over train label counts in each CV for all labels. Ideally they should be almost the same ratio of split of label count between train and test sets.</p>\n\n<p><em>Which is important that the std is very low so that we will not find few labels with larger ratios of train/test than other labels.</em></p>\n\n<p>We should get 28 number for each CV from this and do std and mean checkup on them.</p>\n\n<p><strong>2. Std of CV splits</strong></p>\n\n<p>Here is another useful comparison example can be made on the ratios of train label count of CV[0] over the total train label counts of all CVs. Ideally they should be almost the same. </p>\n\n<p><strong>Which is important that low label categories divided equally between CVs and no single CV get a lot and others get very few labels of these rare categories.</strong></p>\n\n<p>We should get 28 number for each CV from this for train and the same numbers for test, and do std and mean checkup on them </p>\n\n<p><strong>Results</strong>\n1. \nYours: mean, std:   (0.25152254, 0.058061782)\nScikit: mean, std:   (0.25187314, 0.058203448)</p>\n\n<ol>\n<li>\nYours: mean, std:   (0.2, 0.002408708)\nScikit: mean, std:   (0.2, 0.0020237858)</li>\n</ol>\n\n<h2>So it seems both implementations are almost with same performance</h2>\n\n<p><strong>Testing the validity of your approach:</strong>\nTested whether any item in test sets of any of the CV splits is present in the test set of any of the other CV splits.\nResults: False for <em>MultilabelStratifiedKFold</em> and True for <em>MultilabelStratifiedShuffleSplit</em>\nWhich is exactly what is expected!</p>\n\n<p>Code of the test:\n```\nfrom itertools import chain</p>\n\n<p>{True}.issubset(\n    chain.from_iterable(\n        chain.from_iterable(\n            chain.from_iterable(\n                chain.from_iterable(\n                    [[[[[idx in test_indices[n]] \n                        for idx in test_indices[k]] \n                       for n in range(0,len(test_indices)) if n != k ]] \n                     for k in range(0,len(test_indices))]\n                )\n            )\n        )\n    )\n) \n```</p>",
      "rawMarkdown": "It seems Karl used `skmultilearn.model_selection.iterative_train_test_split`  which cannot split into multiple CV splits. it just splits into one train/test set.) **Update:** It seems now this method is `skmultilearn.model_selection.iterative_stratification.iterative_train_test_split(X, y, test_size)`  See the bottom of the page:\n[http://scikit.ml/api/skmultilearn.model_selection.iterative_stratification.html#skmultilearn.model_selection.iterative_stratification.iterative_train_test_split](http://scikit.ml/api/skmultilearn.model_selection.iterative_stratification.html#skmultilearn.model_selection.iterative_stratification.iterative_train_test_split)\n\nI checked that they have CV split method too:\n[http://scikit.ml/api/skmultilearn.model_selection.iterative_stratification.html][1]\n\nBoth methods mentioned in this index page:\nhttp://scikit.ml/genindex.html\n\nI made more extensive std tests with:\n1. Your implementation\nmskf = MultilabelStratifiedKFold(n_splits=5, shuffle=True, random_state=0)\n\n2. Scikit implementation\nmskf = IterativeStratification(n_splits=5, order=1)\n\n**Tests:**\n**1. Std of test/train splits**\n\nHere is a useful comparison example can be made on the ratios of test label count over train label counts in each CV for all labels. Ideally they should be almost the same ratio of split of label count between train and test sets.\n\n*Which is important that the std is very low so that we will not find few labels with larger ratios of train/test than other labels.*\n\nWe should get 28 number for each CV from this and do std and mean checkup on them.\n\n**2. Std of CV splits**\n\nHere is another useful comparison example can be made on the ratios of train label count of CV[0] over the total train label counts of all CVs. Ideally they should be almost the same. \n\n**Which is important that low label categories divided equally between CVs and no single CV get a lot and others get very few labels of these rare categories.**\n\nWe should get 28 number for each CV from this for train and the same numbers for test, and do std and mean checkup on them \n\n**Results**\n1. \nYours: mean, std:   (0.25152254, 0.058061782)\nScikit: mean, std:   (0.25187314, 0.058203448)\n\n2. \nYours: mean, std:   (0.2, 0.002408708)\nScikit: mean, std:   (0.2, 0.0020237858)\n\nSo it seems both implementations are almost with same performance\n----------------------------------------------------\n**Testing the validity of your approach:**\nTested whether any item in test sets of any of the CV splits is present in the test set of any of the other CV splits.\nResults: False for *MultilabelStratifiedKFold* and True for *MultilabelStratifiedShuffleSplit*\nWhich is exactly what is expected!\n\nCode of the test:\n```\nfrom itertools import chain\n\n{True}.issubset(\n    chain.from_iterable(\n        chain.from_iterable(\n            chain.from_iterable(\n                chain.from_iterable(\n                    [[[[[idx in test_indices[n]] \n                        for idx in test_indices[k]] \n                       for n in range(0,len(test_indices)) if n != k ]] \n                     for k in range(0,len(test_indices))]\n                )\n            )\n        )\n    )\n) \n```\n\n\n  [1]: http://scikit.ml/api/skmultilearn.model_selection.iterative_stratification.html",
      "votes": null
    },
    {
      "id": "1015046",
      "postDate": "09/17/2020 22:20:36",
      "content": "<p>I am planning to use this package for the MoA competition. Does iterative-stratification package have an option to maintain balance among label pairs as well?<br>\nThis option is available in scikit-multilearn. The following is taken from the scikit-multilearn docs (<a href=\"http://scikit.ml/stratification.html\" target=\"_blank\">here</a>):<br>\n\"Scikit-multilearn provides an implementation of iterative stratification which aims to provide well-balanced distribution of evidence of label relations up to a <strong>given order</strong>.\"<br>\nIs there a similar option in iterative-stratification package. The problem with scikit-multilearn implementation is that it requires that the target for each sample be a list of labels but the MoA target matrix is an indicator vector/ matrix.</p>\n<p>Thank You.</p>",
      "rawMarkdown": "I am planning to use this package for the MoA competition. Does iterative-stratification package have an option to maintain balance among label pairs as well?\nThis option is available in scikit-multilearn. The following is taken from the scikit-multilearn docs ([here](http://scikit.ml/stratification.html)):\n\"Scikit-multilearn provides an implementation of iterative stratification which aims to provide well-balanced distribution of evidence of label relations up to a **given order**.\"\nIs there a similar option in iterative-stratification package. The problem with scikit-multilearn implementation is that it requires that the target for each sample be a list of labels but the MoA target matrix is an indicator vector/ matrix.\n\nThank You.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1015046,
      "author_name": "para24",
      "author_url": "",
      "post_date": "09/17/2020 22:20:36",
      "content": "<p>I am planning to use this package for the MoA competition. Does iterative-stratification package have an option to maintain balance among label pairs as well?<br>\nThis option is available in scikit-multilearn. The following is taken from the scikit-multilearn docs (<a href=\"http://scikit.ml/stratification.html\" target=\"_blank\">here</a>):<br>\n\"Scikit-multilearn provides an implementation of iterative stratification which aims to provide well-balanced distribution of evidence of label relations up to a <strong>given order</strong>.\"<br>\nIs there a similar option in iterative-stratification package. The problem with scikit-multilearn implementation is that it requires that the target for each sample be a list of labels but the MoA target matrix is an indicator vector/ matrix.</p>\n<p>Thank You.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 399655,
      "author_name": "alber8295",
      "author_url": "",
      "post_date": "10/06/2018 11:57:23",
      "content": "<p>This is really useful for multi label classification. Thank you very much! </p>",
      "votes": null,
      "replies": [
        {
          "id": 399670,
          "author_name": "trentb",
          "author_url": "",
          "post_date": "10/06/2018 13:02:21",
          "content": "<p>You're welcome. Please don't hesitate to let me know of any questions you have or issues you encounter via my GitHub project page. Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 400121,
      "author_name": "kwentar",
      "author_url": "",
      "post_date": "10/07/2018 16:33:59",
      "content": "<p><a href=\"https://www.kaggle.com/kmader/rgb-transfer-learning-with-vgg16-for-protein-atlas\">Here</a> is another way to split data:</p>\n\n<pre><code>raw_train_df, valid_df = train_test_split(image_df, \n                 test_size = 0.3, \n                  # hack to make stratification work                  \n                 stratify = image_df['Target'].map(lambda x: x[:3] if '27' not in x else '0'))\n</code></pre>\n\n<p>Does your approch have advantages over this ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 400153,
          "author_name": "trentb",
          "author_url": "",
          "post_date": "10/07/2018 17:50:53",
          "content": "<p>Good question. It appears that the code you provide performs labelsets-based stratification. This is certainly a valid approach. In <a href=\"http://lpis.csd.auth.gr/publications/sechidis-ecmlpkdd-2011.pdf\">Sechidis et al. (2011)</a>, the authors compare random sampling, labelset-based stratification, and iterative stratification (the method I have implemented) for multiple datasets. There are cases where labelsets-stratification is better than iterative stratification. These cases are typically when the ratio of unique labelsets to the number of samples is small. In cases where the ratio of unique labelsets to the number of samples is not small, iterative stratification is often the better choice.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 400154,
          "author_name": "kwentar",
          "author_url": "",
          "post_date": "10/07/2018 17:53:12",
          "content": "<p>Thank you for answer, will try both of them ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 400549,
          "author_name": "kmader",
          "author_url": "",
          "post_date": "10/08/2018 14:20:48",
          "content": "<p>The approach there (I used it in my notebook) is very much an engineer hack based on gut feelings, not statistical know-how. @Trentb's package looks much better and it would be cool to get this into the standard kaggle kernel toolkit (it would also be useful for the NIH Chest X-Ray data)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 400822,
          "author_name": "trentb",
          "author_url": "",
          "post_date": "10/09/2018 00:28:14",
          "content": "<p>Thank you for the recommendation. I have submitted a pull request to add iterative-stratification to Kaggle Kernels. In the meantime, adding iterative-stratification as a custom package in the settings of a kernel works for me.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417662,
          "author_name": "dawiddabkowski",
          "author_url": "",
          "post_date": "11/08/2018 15:46:57",
          "content": "<p>I do not understand why this 'hack' would work:\n<code>stratify = image_df['Target'].map(lambda x: x[:3] if '27' not in x else '0'))</code>\nIt looks like you just take 3 first characters of each label, I don't know why it would make sense. Could someone please elaborate?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 438702,
          "author_name": "jason2627866800",
          "author_url": "",
          "post_date": "12/14/2018 03:34:10",
          "content": "<p>the map function just split all the multilabel into 74 classes. If don't do this, the classes will be very very large</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 417682,
      "author_name": "ldm314",
      "author_url": "",
      "post_date": "11/08/2018 16:28:56",
      "content": "<p>This looks pretty useful. One questions I have: can it do uneven splits? I don't have the compute power to be running multiple folds, I'd like to do something like train_test_split does where you can specify a test size. I quickly looked through the github and didn't see any options for that.</p>",
      "votes": null,
      "replies": [
        {
          "id": 417864,
          "author_name": "areveillon",
          "author_url": "",
          "post_date": "11/08/2018 22:15:09",
          "content": "<p>if you want for instance 20% test data you can create a 5fold split but only use the first fold. It should ave the good proportion of test / train. If you want 12.5% a 8fold etc...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 418908,
          "author_name": "trentb",
          "author_url": "",
          "post_date": "11/10/2018 21:06:35",
          "content": "<p><a href=\"/ldm314\">@ldm314</a> Yes, you can perform an uneven split. You can call <code>MultilabelStratifiedShuffleSplit(n_splits=1, test_size=0.2, random_state=0)</code> if you want 20% of your data to be for testing and 80% for training.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 418932,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "11/10/2018 22:45:09",
          "content": "<p>Worked great, using it to split the input dataframe. Will see how this splitting method works out, thanks!</p>\n\n<pre>from iterstrat.ml_stratifiers import MultilabelStratifiedShuffleSplit\nimport numpy as np\n\nmsss = MultilabelStratifiedShuffleSplit(n_splits=1, test_size=0.2, random_state=42)\ntrain_df_orig = train_df.copy()\nX = train_df_orig['Id'].tolist()\ny = train_df_orig['target_vec_float'].tolist()\n\nfor train_index, test_index in msss.split(X,y): #it should only do one iteration\n    print(\"TRAIN:\", train_index, \"TEST:\", test_index)\n    train_df = train_df_orig.loc[train_df_orig.index.intersection(train_index)].copy()\n    valid_df = train_df_orig.loc[train_df_orig.index.intersection(test_index)].copy()\ngc.collect()\n</pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 424083,
      "author_name": "hdzheng",
      "author_url": "",
      "post_date": "11/19/2018 14:42:09",
      "content": "<p>I got en error: ValueError: Supported target type is: multilabel-indicator. Got 'binary' instead.  Anyone can help?</p>",
      "votes": null,
      "replies": [
        {
          "id": 424156,
          "author_name": "trentb",
          "author_url": "",
          "post_date": "11/19/2018 16:42:12",
          "content": "<p>Ensure that each target instance is represented as a list/array of the same size and that the lists/arrays consist of 1s and 0s where 1s represent the presence of a subcellular object.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 424381,
          "author_name": "hdzheng",
          "author_url": "",
          "post_date": "11/20/2018 02:36:44",
          "content": "<p>Solved. Thank you very much!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 424318,
      "author_name": "thundo",
      "author_url": "",
      "post_date": "11/19/2018 22:07:49",
      "content": "<p>Don't want to hijack this thread... but can you tell me the difference between this package and <a href=\"http://scikit.ml/\">scikit-multilearn</a> in terms of features?</p>\n\n<p>Thank you</p>",
      "votes": null,
      "replies": [
        {
          "id": 424332,
          "author_name": "trentb",
          "author_url": "",
          "post_date": "11/19/2018 23:12:17",
          "content": "<p>Thundo, I'm glad you brought scikit-multilearn to my attention. I was not aware of another scikit-learn-compatible package that performs multi-label stratification. Looking at the dates in GitHub, I believe multi-label stratification was added to scikit-multilearn around the same time that I discovered the paper on iterative stratification; hence, this is probably why my Google searches didn't find an existing implementation of iterative stratification. Scikit-multilearn offers more than multi-label stratification; whereas, my package solely implements multi-label stratification. I'm interested to hear any comparisons on the two implementations of iterative stratification.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 434742,
          "author_name": "kakadustew",
          "author_url": "",
          "post_date": "12/06/2018 21:53:55",
          "content": "<p>Hi, I have tried both to see what works best. Tested (1)sklearn.model_selection.train_test_split (this did not involve any stratification), (2)skmultilearn.model_selection.iterative_train_test_split, and (3)your implementation. I've also tried to use stratify kwarg for the train_test_split, but found out that it does not work with multilabel examples. I did simple test/train class count ratios on test_size=0.2, then calculated std between the class count ratios of each implementation.\nThis is what I've got:\n(1) std = 0.04475211027333776 (on different rnd state std = 0.04861630879603909\n(2) std = 0.029059979822161946 (on different run I've got 0.06246754109552749 , not sure why, guess it's different rnd state, but there is no option to set the rnd state)\n(3) std = 0.007712835765265891 (changing rnd state produces the same outcome)\nSeems like your technique works best, if a low std is what we want, which I think is the case. \nThe (2) looks pretty broken, but maybe it is designed to handle some stuff which I am not aware of, or it does something desirable what I did not spot. Was looking at the class count distributions, but nothing obvious on why would (2) split it in this uneven way.\nDid not have the balls to check the difference between the (2) and (3) code directly, would definitely clarify stuff, but I think it would give me a proper headache.\nIf I will have some spare time I could try to train on all of those splits to see what gives the best training results.\nFWIW I am rolling with your code, thank you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 434751,
          "author_name": "thundo",
          "author_url": "",
          "post_date": "12/06/2018 22:12:06",
          "content": "<p>About (2) random state... I usually don't set an explicit <code>random_state</code> in the stratificator. However, since sklearn seeds its random state via numpy you can seed everything with <code>np.random.seed(SEED)</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 434821,
          "author_name": "trentb",
          "author_url": "",
          "post_date": "12/07/2018 02:14:51",
          "content": "<p>Karl, thank you for running a comparison test! I'm excited that my code produced such a relatively small std. I'll look into why the rnd state didn't seem to have an effect. In the meantime, something you can try is to set <code>n_splits</code> to 2 (somewhat of a misnomer for scikit-learn's StratifedShuffleSplit) to produce two sets of train/test splits at 80%/20%. They will likely be different as if you had run it twice with <code>n_splits</code> set to 1 and different rnd states.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 435531,
          "author_name": "stecasasso",
          "author_url": "",
          "post_date": "12/08/2018 07:14:49",
          "content": "<p>Hi Karl, <br>\nthanks for running this interesting test! <br>\nDid you or anyone else try different split approaches and reported result in terms of score? (average score of folds). <br>\nThanks in advance</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 458549,
          "author_name": "hwasiti",
          "author_url": "",
          "post_date": "01/19/2019 22:37:13",
          "content": "<p>It seems Karl used <code>skmultilearn.model_selection.iterative_train_test_split</code>  which cannot split into multiple CV splits. it just splits into one train/test set.) <strong>Update:</strong> It seems now this method is <code>skmultilearn.model_selection.iterative_stratification.iterative_train_test_split(X, y, test_size)</code>  See the bottom of the page:\n<a href=\"http://scikit.ml/api/skmultilearn.model_selection.iterative_stratification.html#skmultilearn.model_selection.iterative_stratification.iterative_train_test_split\">http://scikit.ml/api/skmultilearn.model_selection.iterative_stratification.html#skmultilearn.model_selection.iterative_stratification.iterative_train_test_split</a></p>\n\n<p>I checked that they have CV split method too:\n<a href=\"http://scikit.ml/api/skmultilearn.model_selection.iterative_stratification.html\">http://scikit.ml/api/skmultilearn.model_selection.iterative_stratification.html</a></p>\n\n<p>Both methods mentioned in this index page:\n<a href=\"http://scikit.ml/genindex.html\">http://scikit.ml/genindex.html</a></p>\n\n<p>I made more extensive std tests with:\n1. Your implementation\nmskf = MultilabelStratifiedKFold(n_splits=5, shuffle=True, random_state=0)</p>\n\n<ol>\n<li>Scikit implementation\nmskf = IterativeStratification(n_splits=5, order=1)</li>\n</ol>\n\n<p><strong>Tests:</strong>\n*<em>1. Std of test/train splits</em>*</p>\n\n<p>Here is a useful comparison example can be made on the ratios of test label count over train label counts in each CV for all labels. Ideally they should be almost the same ratio of split of label count between train and test sets.</p>\n\n<p><em>Which is important that the std is very low so that we will not find few labels with larger ratios of train/test than other labels.</em></p>\n\n<p>We should get 28 number for each CV from this and do std and mean checkup on them.</p>\n\n<p><strong>2. Std of CV splits</strong></p>\n\n<p>Here is another useful comparison example can be made on the ratios of train label count of CV[0] over the total train label counts of all CVs. Ideally they should be almost the same. </p>\n\n<p><strong>Which is important that low label categories divided equally between CVs and no single CV get a lot and others get very few labels of these rare categories.</strong></p>\n\n<p>We should get 28 number for each CV from this for train and the same numbers for test, and do std and mean checkup on them </p>\n\n<p><strong>Results</strong>\n1. \nYours: mean, std:   (0.25152254, 0.058061782)\nScikit: mean, std:   (0.25187314, 0.058203448)</p>\n\n<ol>\n<li>\nYours: mean, std:   (0.2, 0.002408708)\nScikit: mean, std:   (0.2, 0.0020237858)</li>\n</ol>\n\n<h2>So it seems both implementations are almost with same performance</h2>\n\n<p><strong>Testing the validity of your approach:</strong>\nTested whether any item in test sets of any of the CV splits is present in the test set of any of the other CV splits.\nResults: False for <em>MultilabelStratifiedKFold</em> and True for <em>MultilabelStratifiedShuffleSplit</em>\nWhich is exactly what is expected!</p>\n\n<p>Code of the test:\n```\nfrom itertools import chain</p>\n\n<p>{True}.issubset(\n    chain.from_iterable(\n        chain.from_iterable(\n            chain.from_iterable(\n                chain.from_iterable(\n                    [[[[[idx in test_indices[n]] \n                        for idx in test_indices[k]] \n                       for n in range(0,len(test_indices)) if n != k ]] \n                     for k in range(0,len(test_indices))]\n                )\n            )\n        )\n    )\n) \n```</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 434812,
      "author_name": "marsyeti",
      "author_url": "",
      "post_date": "12/07/2018 01:34:28",
      "content": "<p>Thank you very much for your suggestion! I don't have too much time to commit to this competition, but I also would like to share my idea. \nMy data loader just iterate through each class indices and feed them evenly into my model(e.g. 2 per class, makes batch size 56), when one class indices run out, just reshuffle it. This requires more data augmentation than usual imo, but not sure.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 434908,
      "author_name": "irvingzhang0512",
      "author_url": "",
      "post_date": "12/07/2018 06:13:43",
      "content": "<p>thanks a lot, very helpful. im using MultilabelStratifiedShuffleSplit now.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "399431": "Participants may find it helpful to balance distributions of multilabel data across splits for cross validation (i.e., stratify the data). Earlier this year I created a Python package called iterative-stratification that aims to accomplish this task for multilabel data: https://github.com/trent-b/iterative-stratification. I hope the package may find some utility in this competition.",
    "399655": "This is really useful for multi label classification. Thank you very much!",
    "399670": "You're welcome. Please don't hesitate to let me know of any questions you have or issues you encounter via my GitHub project page. Thanks!",
    "400121": "[Here][1] is another way to split data:\n\n    raw_train_df, valid_df = train_test_split(image_df, \n                     test_size = 0.3, \n                      # hack to make stratification work                  \n                     stratify = image_df['Target'].map(lambda x: x[:3] if '27' not in x else '0'))\n\nDoes your approch have advantages over this ? \n\n  [1]: https://www.kaggle.com/kmader/rgb-transfer-learning-with-vgg16-for-protein-atlas",
    "400153": "Good question. It appears that the code you provide performs labelsets-based stratification. This is certainly a valid approach. In [Sechidis et al. (2011)][1], the authors compare random sampling, labelset-based stratification, and iterative stratification (the method I have implemented) for multiple datasets. There are cases where labelsets-stratification is better than iterative stratification. These cases are typically when the ratio of unique labelsets to the number of samples is small. In cases where the ratio of unique labelsets to the number of samples is not small, iterative stratification is often the better choice.\n\n\n  [1]: http://lpis.csd.auth.gr/publications/sechidis-ecmlpkdd-2011.pdf",
    "400154": "Thank you for answer, will try both of them ;)",
    "400549": "The approach there (I used it in my notebook) is very much an engineer hack based on gut feelings, not statistical know-how. @Trentb's package looks much better and it would be cool to get this into the standard kaggle kernel toolkit (it would also be useful for the NIH Chest X-Ray data)",
    "400822": "Thank you for the recommendation. I have submitted a pull request to add iterative-stratification to Kaggle Kernels. In the meantime, adding iterative-stratification as a custom package in the settings of a kernel works for me.",
    "417662": "I do not understand why this 'hack' would work:\n`stratify = image_df['Target'].map(lambda x: x[:3] if '27' not in x else '0'))`\nIt looks like you just take 3 first characters of each label, I don't know why it would make sense. Could someone please elaborate?",
    "417682": "This looks pretty useful. One questions I have: can it do uneven splits? I don't have the compute power to be running multiple folds, I'd like to do something like train_test_split does where you can specify a test size. I quickly looked through the github and didn't see any options for that.",
    "417864": "if you want for instance 20% test data you can create a 5fold split but only use the first fold. It should ave the good proportion of test / train. If you want 12.5% a 8fold etc...",
    "418908": "ldm314 Yes, you can perform an uneven split. You can call `MultilabelStratifiedShuffleSplit(n_splits=1, test_size=0.2, random_state=0)` if you want 20% of your data to be for testing and 80% for training.",
    "418932": "Worked great, using it to split the input dataframe. Will see how this splitting method works out, thanks!\n\n<pre>from iterstrat.ml_stratifiers import MultilabelStratifiedShuffleSplit\nimport numpy as np\n\nmsss = MultilabelStratifiedShuffleSplit(n_splits=1, test_size=0.2, random_state=42)\ntrain_df_orig = train_df.copy()\nX = train_df_orig['Id'].tolist()\ny = train_df_orig['target_vec_float'].tolist()\n\nfor train_index, test_index in msss.split(X,y): #it should only do one iteration\n    print(\"TRAIN:\", train_index, \"TEST:\", test_index)\n    train_df = train_df_orig.loc[train_df_orig.index.intersection(train_index)].copy()\n    valid_df = train_df_orig.loc[train_df_orig.index.intersection(test_index)].copy()\ngc.collect()\n</pre>",
    "424083": "I got en error: ValueError: Supported target type is: multilabel-indicator. Got 'binary' instead.  Anyone can help?",
    "424156": "Ensure that each target instance is represented as a list/array of the same size and that the lists/arrays consist of 1s and 0s where 1s represent the presence of a subcellular object.",
    "424318": "Don't want to hijack this thread... but can you tell me the difference between this package and [scikit-multilearn](http://scikit.ml/) in terms of features?\n\nThank you",
    "424332": "Thundo, I'm glad you brought scikit-multilearn to my attention. I was not aware of another scikit-learn-compatible package that performs multi-label stratification. Looking at the dates in GitHub, I believe multi-label stratification was added to scikit-multilearn around the same time that I discovered the paper on iterative stratification; hence, this is probably why my Google searches didn't find an existing implementation of iterative stratification. Scikit-multilearn offers more than multi-label stratification; whereas, my package solely implements multi-label stratification. I'm interested to hear any comparisons on the two implementations of iterative stratification.",
    "424381": "Solved. Thank you very much!",
    "434742": "Hi, I have tried both to see what works best. Tested (1)sklearn.model_selection.train_test_split (this did not involve any stratification), (2)skmultilearn.model_selection.iterative_train_test_split, and (3)your implementation. I've also tried to use stratify kwarg for the train_test_split, but found out that it does not work with multilabel examples. I did simple test/train class count ratios on test_size=0.2, then calculated std between the class count ratios of each implementation.\nThis is what I've got:\n(1) std = 0.04475211027333776 (on different rnd state std = 0.04861630879603909\n(2) std = 0.029059979822161946 (on different run I've got 0.06246754109552749 , not sure why, guess it's different rnd state, but there is no option to set the rnd state)\n(3) std = 0.007712835765265891 (changing rnd state produces the same outcome)\nSeems like your technique works best, if a low std is what we want, which I think is the case. \nThe (2) looks pretty broken, but maybe it is designed to handle some stuff which I am not aware of, or it does something desirable what I did not spot. Was looking at the class count distributions, but nothing obvious on why would (2) split it in this uneven way.\nDid not have the balls to check the difference between the (2) and (3) code directly, would definitely clarify stuff, but I think it would give me a proper headache.\nIf I will have some spare time I could try to train on all of those splits to see what gives the best training results.\nFWIW I am rolling with your code, thank you.",
    "434751": "About (2) random state... I usually don't set an explicit `random_state` in the stratificator. However, since sklearn seeds its random state via numpy you can seed everything with `np.random.seed(SEED)`",
    "434812": "Thank you very much for your suggestion! I don't have too much time to commit to this competition, but I also would like to share my idea. \nMy data loader just iterate through each class indices and feed them evenly into my model(e.g. 2 per class, makes batch size 56), when one class indices run out, just reshuffle it. This requires more data augmentation than usual imo, but not sure.",
    "434821": "Karl, thank you for running a comparison test! I'm excited that my code produced such a relatively small std. I'll look into why the rnd state didn't seem to have an effect. In the meantime, something you can try is to set ```n_splits``` to 2 (somewhat of a misnomer for scikit-learn's StratifedShuffleSplit) to produce two sets of train/test splits at 80%/20%. They will likely be different as if you had run it twice with ```n_splits``` set to 1 and different rnd states.",
    "434908": "thanks a lot, very helpful. im using MultilabelStratifiedShuffleSplit now.",
    "435531": "Hi Karl,   \nthanks for running this interesting test!  \nDid you or anyone else try different split approaches and reported result in terms of score? (average score of folds).   \nThanks in advance",
    "438702": "the map function just split all the multilabel into 74 classes. If don't do this, the classes will be very very large",
    "458549": "It seems Karl used `skmultilearn.model_selection.iterative_train_test_split`  which cannot split into multiple CV splits. it just splits into one train/test set.) **Update:** It seems now this method is `skmultilearn.model_selection.iterative_stratification.iterative_train_test_split(X, y, test_size)`  See the bottom of the page:\n[http://scikit.ml/api/skmultilearn.model_selection.iterative_stratification.html#skmultilearn.model_selection.iterative_stratification.iterative_train_test_split](http://scikit.ml/api/skmultilearn.model_selection.iterative_stratification.html#skmultilearn.model_selection.iterative_stratification.iterative_train_test_split)\n\nI checked that they have CV split method too:\n[http://scikit.ml/api/skmultilearn.model_selection.iterative_stratification.html][1]\n\nBoth methods mentioned in this index page:\nhttp://scikit.ml/genindex.html\n\nI made more extensive std tests with:\n1. Your implementation\nmskf = MultilabelStratifiedKFold(n_splits=5, shuffle=True, random_state=0)\n\n2. Scikit implementation\nmskf = IterativeStratification(n_splits=5, order=1)\n\n**Tests:**\n**1. Std of test/train splits**\n\nHere is a useful comparison example can be made on the ratios of test label count over train label counts in each CV for all labels. Ideally they should be almost the same ratio of split of label count between train and test sets.\n\n*Which is important that the std is very low so that we will not find few labels with larger ratios of train/test than other labels.*\n\nWe should get 28 number for each CV from this and do std and mean checkup on them.\n\n**2. Std of CV splits**\n\nHere is another useful comparison example can be made on the ratios of train label count of CV[0] over the total train label counts of all CVs. Ideally they should be almost the same. \n\n**Which is important that low label categories divided equally between CVs and no single CV get a lot and others get very few labels of these rare categories.**\n\nWe should get 28 number for each CV from this for train and the same numbers for test, and do std and mean checkup on them \n\n**Results**\n1. \nYours: mean, std:   (0.25152254, 0.058061782)\nScikit: mean, std:   (0.25187314, 0.058203448)\n\n2. \nYours: mean, std:   (0.2, 0.002408708)\nScikit: mean, std:   (0.2, 0.0020237858)\n\nSo it seems both implementations are almost with same performance\n----------------------------------------------------\n**Testing the validity of your approach:**\nTested whether any item in test sets of any of the CV splits is present in the test set of any of the other CV splits.\nResults: False for *MultilabelStratifiedKFold* and True for *MultilabelStratifiedShuffleSplit*\nWhich is exactly what is expected!\n\nCode of the test:\n```\nfrom itertools import chain\n\n{True}.issubset(\n    chain.from_iterable(\n        chain.from_iterable(\n            chain.from_iterable(\n                chain.from_iterable(\n                    [[[[[idx in test_indices[n]] \n                        for idx in test_indices[k]] \n                       for n in range(0,len(test_indices)) if n != k ]] \n                     for k in range(0,len(test_indices))]\n                )\n            )\n        )\n    )\n) \n```\n\n\n  [1]: http://scikit.ml/api/skmultilearn.model_selection.iterative_stratification.html",
    "1015046": "I am planning to use this package for the MoA competition. Does iterative-stratification package have an option to maintain balance among label pairs as well?\nThis option is available in scikit-multilearn. The following is taken from the scikit-multilearn docs ([here](http://scikit.ml/stratification.html)):\n\"Scikit-multilearn provides an implementation of iterative stratification which aims to provide well-balanced distribution of evidence of label relations up to a **given order**.\"\nIs there a similar option in iterative-stratification package. The problem with scikit-multilearn implementation is that it requires that the target for each sample be a list of labels but the MoA target matrix is an indicator vector/ matrix.\n\nThank You."
  },
  "source": "meta"
}