{
  "id": 157973,
  "title": "Combining Tabular + Vision in fastai2",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/157973",
  "author_name": "Zach Mueller",
  "post_date": "2020-06-12T19:00:07.511000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi everyone, I've published a Kernel discussing how to do such an integration in the fastai2 library. In the previous version, this was <em>super</em> tiresome to do and didn't provide much wiggleroom, especially for image/data specific augmentation. The kernel (<a href=\"https://www.kaggle.com/muellerzr/fastai2-tabular-vision-starter-kernel\">here</a>) describes the new process. It can be summed up as we make a <code>DataLoader</code> of <code>DataLoaders</code>. <code>fastai2</code> calls the shuffled batches via index's, so we just need to ensure those index's always match up. This can very easily be expanded to not only <code>n</code> <code>DataLoaders</code> but also can work with <em>any</em> <code>DataLoader</code> itself, as you're pulling and merging the raw transformed batches <em>inside</em> the fastai pipeline. If there are any comments or questions about the process, feel free to ask here!</p>\n\n<p>Here's the idea (for this competition specifically):</p>\n\n<h2>The Pipeline</h2>\n\n<p>Here is an outline of how you go about doing this:\n1. Make your <code>tab</code> and <code>vis</code> <code>DataLoaders</code>\n    (<code>vis</code> = Vision, <code>tab</code> = Tabular)\n2. Combine them together into a <code>Hybrid DataLoader</code>\n3. Adjust your own <code>test_dl</code> framework how you choose\n4. Train</p>\n\n<h2>The Code:</h2>\n\n<p>Now let's talk about the code. For our \"DataLoader\", it won't inherit the <code>DataLoader</code> class (hence the quotes around it). Instead we'll give it the <em>minimal similar behavior</em> to a <code>DataLoader</code> that is needed, and have everything else work internally. Specifically, these functionalities:\n* <code>FakeLoader</code>\n* <code>__len__</code>\n* <code>__iter__</code>\n* <code>one_batch</code>\n* <code>show_batch</code>\n* <code>shuffle_fn</code>\n* <code>to</code></p>\n\n<p>Now to build this I'm going to walk us through it with <code>@patch</code> from the <code>fastcore</code> library. Basically this lets us lazily define the class as we go, so don't get confused to why it's all in more than one block.</p>\n\n<h2><code>__init__</code> and <code>FakeLoader</code></h2>\n\n<p>The <code>__init__</code> for our model needs to store 5 items, the <code>device</code> we're running on, our two <code>DataLoaders</code> we're passing in, a <code>count</code>, a <code>_FakeLoader</code>, and our new shuffle function (for now this will be undefined, we'll discuss it more in a moment). Also, <code>FakeLoader</code> is used during the <code>__iter__</code>, see the regular <code>DataLoader</code> source code to see it there:</p>\n\n<p><code>python\nfrom fastai2.data.load import _FakeLoader, _loaders\nclass MixedDL():\n    def __init__(self, tab_dl:TabDataLoader, vis_dl:TfmdDL, device='cuda:0'):\n        \"Stores away `tab_dl` and `vis_dl`, and overrides `shuffle_fn`\"\n        self.device = device\n        tab_dl.shuffle_fn = self.shuffle_fn\n        vis_dl.shuffle_fn = self.shuffle_fn\n        self.dls = [tab_dl, vis_dl]\n        self.count = 0\n        self.fake_l = _FakeLoader(self, False, 0, 0)\n</code></p>\n\n<h2><code>shuffle_fn</code></h2>\n\n<p>Now we'll look at the <code>shuffle_fn</code> there. What needs to have happen? The <code>shuffle_fn</code> returns a list of index's for us to use, that's stored inside of <code>self.rng</code>, and we want those index's to change every 2 times we call the <code>shuffle_fn</code> (as we call it for each of our internal <code>DataLoaders</code>), to ensure that both are mapped out to the same index's for preparing our batch. This is what that looks like:\n<code>python\n@patch <br>\ndef shuffle_fn(x:MixedDL, idxs):\n        \"Generates a new `rng` based upon which `DataLoader` is called\"\n        if self.count == 0: # if we haven't generated an rng yet\n            self.rng = self.dls[0].rng.sample(idxs, len(idxs))\n            self.count += 1\n            return self.rng\n        else:\n            self.count = 0\n            return self.rng\n</code>\nThis is <strong>all</strong> that's needed to ensure that all of our batches get shuffled together. And if you're using more than two, count is just equal to <code>n</code> internal <code>DataLoaders</code>.</p>\n\n<p>While we're at it, we'll take care of two other functions, the <code>__len__</code> attribute and the <code>to</code> function. <code>__len__</code> just needs to grab the length of <em>one</em> of our <code>DataLoaders</code>, and <code>to</code> just returns the name of our device:</p>\n\n<p>```python\n@patch \ndef <strong>len</strong>(x:HybridDL): return len(x.dls[0])</p>\n\n<p>@patch\ndef to(x:HybridDL, device): x.device = device\n```</p>\n\n<h2><code>__iter__</code></h2>\n\n<p>Now let's move into something a bit more complex, the iterator. Now, our iterator needs to take all of our batches from our loaders and perform the <code>after_batch</code> transform for those outputs from <em>their respective <code>DataLoader</code></em> before finally being put into a batch, also moving each to the <code>device</code>. While this may look scary, the <code>_loaders</code> etc is all the same as it is from the <code>DataLoaders</code> class, so it's just how we access them:</p>\n\n<p><code>python\n@patch\ndef __iter__(dl:MixedDL):\n    \"Iterate over your `DataLoader`\"\n    z = zip(*[_loaders[i.fake_l.num_workers==0](i.fake_l) for i in dl.dls])\n    for b in z:\n        if dl.device is not None: \n            b = to_device(b, dl.device)\n        batch = []\n        batch.extend(dl.dls[0].after_batch(b[0])[:2]) # tabular cat and cont\n        batch.append(dl.dls[1].after_batch(b[1][0])) # Image\n        try: # In case the data is unlabelled\n            batch.append(b[1][1]) # y\n            yield tuple(batch)\n        except:\n            yield tuple(batch)\n</code>\nNotice the device is adjusted recursively before we move to the batch transforms (this is how <code>fastai</code> moves them all to the GPU)</p>\n\n<h2><code>one_batch</code></h2>\n\n<p>Alright, so we can build it, iterate it, now how do we get our good ol' fashion <code>one_batch</code>? Quite easily. We call <code>fake_l.no_multiproc()</code> (which so you know, that means we temporarily adjust the <code>num_workers</code> in our <code>DataLoader</code> to zero) and grab the first batch, while also discarding any iterators the <code>DataLoader</code> may have (as <code>first</code> calls <code>next(iter(dl))</code>):\n<code>python\n@patch\ndef one_batch(x:MixedDL):\n    \"Grab a batch from the `DataLoader`\"\n    with x.fake_l.no_multiproc(): res = first(x)\n    if hasattr(x, 'it'): delattr(x, 'it')\n    return res\n</code>\nYou <em>may</em> or may not get an exception error, this can be safely ignored. Your batch now returns as <code>[cat, cont, im, y]</code></p>\n\n<h2><code>show_batch</code></h2>\n\n<p>Next up is probably the easiest out of all of the functions. All we're wanting to do here is in each <code>DataLoader</code>, call <code>show_batch</code>. It's as simple as it sounds:\n<code>python\n@patch\ndef show_batch(x:MixedDL):\n    \"Show a batch from multiple `DataLoaders`\"\n    for dl in x.dls:\n        dl.show_batch()\n</code>\nHere's an example output:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/883597/15850/26cee5dec222e423748e26b14ae2776011974ae6.jpeg\" alt=\"\"></p>\n\n<p>And that's <strong>all</strong> that's needed to start training and have <em>all</em> the functionalities of <code>fastai</code> while bringing in the various <code>DataTypes</code>. So they key that made this entire thing possible is due to how <code>fastai</code> does the <code>shuffle_fn</code>, and the fact they are indices. </p>\n\n<h2><code>test_dl</code></h2>\n\n<p>The last thing I'll show is how to do the <code>test_dl</code>. When you're making these ideally you build the entire Image and Tabular <code>dls</code>, which gives you access to the <code>.test_dl</code> function. From there, simply do something like:\n<code>python\nim_test = vis_dl.test_dl(test_df)\ntab_test = tab_dl.test_dl(test_df)\ntest_dl = MixedDL(tab_test, im_test)\n</code>\nAnd you're good to go! The main reason we don't have to worry about enabling shuffling, etc is due to the fact it's done on the <em>interior</em> <code>DataLoader</code> level.</p>\n\n<p>Let me know any questions you have!</p>",
  "messages": [
    {
      "id": 883597,
      "postDate": "2020-06-12T19:00:07.513Z",
      "content": "<p>Hi everyone, I've published a Kernel discussing how to do such an integration in the fastai2 library. In the previous version, this was <em>super</em> tiresome to do and didn't provide much wiggleroom, especially for image/data specific augmentation. The kernel (<a href=\"https://www.kaggle.com/muellerzr/fastai2-tabular-vision-starter-kernel\">here</a>) describes the new process. It can be summed up as we make a <code>DataLoader</code> of <code>DataLoaders</code>. <code>fastai2</code> calls the shuffled batches via index's, so we just need to ensure those index's always match up. This can very easily be expanded to not only <code>n</code> <code>DataLoaders</code> but also can work with <em>any</em> <code>DataLoader</code> itself, as you're pulling and merging the raw transformed batches <em>inside</em> the fastai pipeline. If there are any comments or questions about the process, feel free to ask here!</p>\n\n<p>Here's the idea (for this competition specifically):</p>\n\n<h2>The Pipeline</h2>\n\n<p>Here is an outline of how you go about doing this:\n1. Make your <code>tab</code> and <code>vis</code> <code>DataLoaders</code>\n    (<code>vis</code> = Vision, <code>tab</code> = Tabular)\n2. Combine them together into a <code>Hybrid DataLoader</code>\n3. Adjust your own <code>test_dl</code> framework how you choose\n4. Train</p>\n\n<h2>The Code:</h2>\n\n<p>Now let's talk about the code. For our \"DataLoader\", it won't inherit the <code>DataLoader</code> class (hence the quotes around it). Instead we'll give it the <em>minimal similar behavior</em> to a <code>DataLoader</code> that is needed, and have everything else work internally. Specifically, these functionalities:\n* <code>FakeLoader</code>\n* <code>__len__</code>\n* <code>__iter__</code>\n* <code>one_batch</code>\n* <code>show_batch</code>\n* <code>shuffle_fn</code>\n* <code>to</code></p>\n\n<p>Now to build this I'm going to walk us through it with <code>@patch</code> from the <code>fastcore</code> library. Basically this lets us lazily define the class as we go, so don't get confused to why it's all in more than one block.</p>\n\n<h2><code>__init__</code> and <code>FakeLoader</code></h2>\n\n<p>The <code>__init__</code> for our model needs to store 5 items, the <code>device</code> we're running on, our two <code>DataLoaders</code> we're passing in, a <code>count</code>, a <code>_FakeLoader</code>, and our new shuffle function (for now this will be undefined, we'll discuss it more in a moment). Also, <code>FakeLoader</code> is used during the <code>__iter__</code>, see the regular <code>DataLoader</code> source code to see it there:</p>\n\n<p><code>python\nfrom fastai2.data.load import _FakeLoader, _loaders\nclass MixedDL():\n    def __init__(self, tab_dl:TabDataLoader, vis_dl:TfmdDL, device='cuda:0'):\n        \"Stores away `tab_dl` and `vis_dl`, and overrides `shuffle_fn`\"\n        self.device = device\n        tab_dl.shuffle_fn = self.shuffle_fn\n        vis_dl.shuffle_fn = self.shuffle_fn\n        self.dls = [tab_dl, vis_dl]\n        self.count = 0\n        self.fake_l = _FakeLoader(self, False, 0, 0)\n</code></p>\n\n<h2><code>shuffle_fn</code></h2>\n\n<p>Now we'll look at the <code>shuffle_fn</code> there. What needs to have happen? The <code>shuffle_fn</code> returns a list of index's for us to use, that's stored inside of <code>self.rng</code>, and we want those index's to change every 2 times we call the <code>shuffle_fn</code> (as we call it for each of our internal <code>DataLoaders</code>), to ensure that both are mapped out to the same index's for preparing our batch. This is what that looks like:\n<code>python\n@patch <br>\ndef shuffle_fn(x:MixedDL, idxs):\n        \"Generates a new `rng` based upon which `DataLoader` is called\"\n        if self.count == 0: # if we haven't generated an rng yet\n            self.rng = self.dls[0].rng.sample(idxs, len(idxs))\n            self.count += 1\n            return self.rng\n        else:\n            self.count = 0\n            return self.rng\n</code>\nThis is <strong>all</strong> that's needed to ensure that all of our batches get shuffled together. And if you're using more than two, count is just equal to <code>n</code> internal <code>DataLoaders</code>.</p>\n\n<p>While we're at it, we'll take care of two other functions, the <code>__len__</code> attribute and the <code>to</code> function. <code>__len__</code> just needs to grab the length of <em>one</em> of our <code>DataLoaders</code>, and <code>to</code> just returns the name of our device:</p>\n\n<p>```python\n@patch \ndef <strong>len</strong>(x:HybridDL): return len(x.dls[0])</p>\n\n<p>@patch\ndef to(x:HybridDL, device): x.device = device\n```</p>\n\n<h2><code>__iter__</code></h2>\n\n<p>Now let's move into something a bit more complex, the iterator. Now, our iterator needs to take all of our batches from our loaders and perform the <code>after_batch</code> transform for those outputs from <em>their respective <code>DataLoader</code></em> before finally being put into a batch, also moving each to the <code>device</code>. While this may look scary, the <code>_loaders</code> etc is all the same as it is from the <code>DataLoaders</code> class, so it's just how we access them:</p>\n\n<p><code>python\n@patch\ndef __iter__(dl:MixedDL):\n    \"Iterate over your `DataLoader`\"\n    z = zip(*[_loaders[i.fake_l.num_workers==0](i.fake_l) for i in dl.dls])\n    for b in z:\n        if dl.device is not None: \n            b = to_device(b, dl.device)\n        batch = []\n        batch.extend(dl.dls[0].after_batch(b[0])[:2]) # tabular cat and cont\n        batch.append(dl.dls[1].after_batch(b[1][0])) # Image\n        try: # In case the data is unlabelled\n            batch.append(b[1][1]) # y\n            yield tuple(batch)\n        except:\n            yield tuple(batch)\n</code>\nNotice the device is adjusted recursively before we move to the batch transforms (this is how <code>fastai</code> moves them all to the GPU)</p>\n\n<h2><code>one_batch</code></h2>\n\n<p>Alright, so we can build it, iterate it, now how do we get our good ol' fashion <code>one_batch</code>? Quite easily. We call <code>fake_l.no_multiproc()</code> (which so you know, that means we temporarily adjust the <code>num_workers</code> in our <code>DataLoader</code> to zero) and grab the first batch, while also discarding any iterators the <code>DataLoader</code> may have (as <code>first</code> calls <code>next(iter(dl))</code>):\n<code>python\n@patch\ndef one_batch(x:MixedDL):\n    \"Grab a batch from the `DataLoader`\"\n    with x.fake_l.no_multiproc(): res = first(x)\n    if hasattr(x, 'it'): delattr(x, 'it')\n    return res\n</code>\nYou <em>may</em> or may not get an exception error, this can be safely ignored. Your batch now returns as <code>[cat, cont, im, y]</code></p>\n\n<h2><code>show_batch</code></h2>\n\n<p>Next up is probably the easiest out of all of the functions. All we're wanting to do here is in each <code>DataLoader</code>, call <code>show_batch</code>. It's as simple as it sounds:\n<code>python\n@patch\ndef show_batch(x:MixedDL):\n    \"Show a batch from multiple `DataLoaders`\"\n    for dl in x.dls:\n        dl.show_batch()\n</code>\nHere's an example output:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/883597/15850/26cee5dec222e423748e26b14ae2776011974ae6.jpeg\" alt=\"\"></p>\n\n<p>And that's <strong>all</strong> that's needed to start training and have <em>all</em> the functionalities of <code>fastai</code> while bringing in the various <code>DataTypes</code>. So they key that made this entire thing possible is due to how <code>fastai</code> does the <code>shuffle_fn</code>, and the fact they are indices. </p>\n\n<h2><code>test_dl</code></h2>\n\n<p>The last thing I'll show is how to do the <code>test_dl</code>. When you're making these ideally you build the entire Image and Tabular <code>dls</code>, which gives you access to the <code>.test_dl</code> function. From there, simply do something like:\n<code>python\nim_test = vis_dl.test_dl(test_df)\ntab_test = tab_dl.test_dl(test_df)\ntest_dl = MixedDL(tab_test, im_test)\n</code>\nAnd you're good to go! The main reason we don't have to worry about enabling shuffling, etc is due to the fact it's done on the <em>interior</em> <code>DataLoader</code> level.</p>\n\n<p>Let me know any questions you have!</p>",
      "rawMarkdown": "Hi everyone, I've published a Kernel discussing how to do such an integration in the fastai2 library. In the previous version, this was *super* tiresome to do and didn't provide much wiggleroom, especially for image/data specific augmentation. The kernel ([here](https://www.kaggle.com/muellerzr/fastai2-tabular-vision-starter-kernel)) describes the new process. It can be summed up as we make a `DataLoader` of `DataLoaders`. `fastai2` calls the shuffled batches via index's, so we just need to ensure those index's always match up. This can very easily be expanded to not only `n` `DataLoaders` but also can work with *any* `DataLoader` itself, as you're pulling and merging the raw transformed batches *inside* the fastai pipeline. If there are any comments or questions about the process, feel free to ask here!\n\nHere's the idea (for this competition specifically):\n\n## The Pipeline\nHere is an outline of how you go about doing this:\n1. Make your `tab` and `vis` `DataLoaders`\n    (`vis` = Vision, `tab` = Tabular)\n2. Combine them together into a `Hybrid DataLoader`\n3. Adjust your own `test_dl` framework how you choose\n4. Train\n\n## The Code:\nNow let's talk about the code. For our \"DataLoader\", it won't inherit the `DataLoader` class (hence the quotes around it). Instead we'll give it the *minimal similar behavior* to a `DataLoader` that is needed, and have everything else work internally. Specifically, these functionalities:\n* `FakeLoader`\n* `__len__`\n* `__iter__`\n* `one_batch`\n* `show_batch`\n* `shuffle_fn`\n* `to`\n\nNow to build this I'm going to walk us through it with `@patch` from the `fastcore` library. Basically this lets us lazily define the class as we go, so don't get confused to why it's all in more than one block.\n\n## `__init__` and `FakeLoader`\nThe `__init__` for our model needs to store 5 items, the `device` we're running on, our two `DataLoaders` we're passing in, a `count`, a `_FakeLoader`, and our new shuffle function (for now this will be undefined, we'll discuss it more in a moment). Also, `FakeLoader` is used during the `__iter__`, see the regular `DataLoader` source code to see it there:\n\n```python\nfrom fastai2.data.load import _FakeLoader, _loaders\nclass MixedDL():\n    def __init__(self, tab_dl:TabDataLoader, vis_dl:TfmdDL, device='cuda:0'):\n        \"Stores away `tab_dl` and `vis_dl`, and overrides `shuffle_fn`\"\n        self.device = device\n        tab_dl.shuffle_fn = self.shuffle_fn\n        vis_dl.shuffle_fn = self.shuffle_fn\n        self.dls = [tab_dl, vis_dl]\n        self.count = 0\n        self.fake_l = _FakeLoader(self, False, 0, 0)\n```\n\n## `shuffle_fn`\nNow we'll look at the `shuffle_fn` there. What needs to have happen? The `shuffle_fn` returns a list of index's for us to use, that's stored inside of `self.rng`, and we want those index's to change every 2 times we call the `shuffle_fn` (as we call it for each of our internal `DataLoaders`), to ensure that both are mapped out to the same index's for preparing our batch. This is what that looks like:\n```python\n@patch   \ndef shuffle_fn(x:MixedDL, idxs):\n        \"Generates a new `rng` based upon which `DataLoader` is called\"\n        if self.count == 0: # if we haven't generated an rng yet\n            self.rng = self.dls[0].rng.sample(idxs, len(idxs))\n            self.count += 1\n            return self.rng\n        else:\n            self.count = 0\n            return self.rng\n```\nThis is **all** that's needed to ensure that all of our batches get shuffled together. And if you're using more than two, count is just equal to `n` internal `DataLoaders`.\n\nWhile we're at it, we'll take care of two other functions, the `__len__` attribute and the `to` function. `__len__` just needs to grab the length of *one* of our `DataLoaders`, and `to` just returns the name of our device:\n\n```python\n@patch \ndef __len__(x:HybridDL): return len(x.dls[0])\n\n@patch\ndef to(x:HybridDL, device): x.device = device\n```\n\n## `__iter__`\nNow let's move into something a bit more complex, the iterator. Now, our iterator needs to take all of our batches from our loaders and perform the `after_batch` transform for those outputs from *their respective `DataLoader`* before finally being put into a batch, also moving each to the `device`. While this may look scary, the `_loaders` etc is all the same as it is from the `DataLoaders` class, so it's just how we access them:\n\n```python\n@patch\ndef __iter__(dl:MixedDL):\n    \"Iterate over your `DataLoader`\"\n    z = zip(*[_loaders[i.fake_l.num_workers==0](i.fake_l) for i in dl.dls])\n    for b in z:\n        if dl.device is not None: \n            b = to_device(b, dl.device)\n        batch = []\n        batch.extend(dl.dls[0].after_batch(b[0])[:2]) # tabular cat and cont\n        batch.append(dl.dls[1].after_batch(b[1][0])) # Image\n        try: # In case the data is unlabelled\n            batch.append(b[1][1]) # y\n            yield tuple(batch)\n        except:\n            yield tuple(batch)\n```\nNotice the device is adjusted recursively before we move to the batch transforms (this is how `fastai` moves them all to the GPU)\n\n## `one_batch`\n\nAlright, so we can build it, iterate it, now how do we get our good ol' fashion `one_batch`? Quite easily. We call `fake_l.no_multiproc()` (which so you know, that means we temporarily adjust the `num_workers` in our `DataLoader` to zero) and grab the first batch, while also discarding any iterators the `DataLoader` may have (as `first` calls `next(iter(dl))`):\n```python\n@patch\ndef one_batch(x:MixedDL):\n    \"Grab a batch from the `DataLoader`\"\n    with x.fake_l.no_multiproc(): res = first(x)\n    if hasattr(x, 'it'): delattr(x, 'it')\n    return res\n```\nYou *may* or may not get an exception error, this can be safely ignored. Your batch now returns as `[cat, cont, im, y]`\n\n## `show_batch`\n\nNext up is probably the easiest out of all of the functions. All we're wanting to do here is in each `DataLoader`, call `show_batch`. It's as simple as it sounds:\n```python\n@patch\ndef show_batch(x:MixedDL):\n    \"Show a batch from multiple `DataLoaders`\"\n    for dl in x.dls:\n        dl.show_batch()\n```\nHere's an example output:\n\n![](https://storage.googleapis.com/kaggle-forum-message-attachments/883597/15850/26cee5dec222e423748e26b14ae2776011974ae6.jpeg)\n\nAnd that's **all** that's needed to start training and have *all* the functionalities of `fastai` while bringing in the various `DataTypes`. So they key that made this entire thing possible is due to how `fastai` does the `shuffle_fn`, and the fact they are indices. \n\n## `test_dl`\n\nThe last thing I'll show is how to do the `test_dl`. When you're making these ideally you build the entire Image and Tabular `dls`, which gives you access to the `.test_dl` function. From there, simply do something like:\n```python\nim_test = vis_dl.test_dl(test_df)\ntab_test = tab_dl.test_dl(test_df)\ntest_dl = MixedDL(tab_test, im_test)\n```\nAnd you're good to go! The main reason we don't have to worry about enabling shuffling, etc is due to the fact it's done on the *interior* `DataLoader` level.\n\nLet me know any questions you have!",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "883597": "Hi everyone, I've published a Kernel discussing how to do such an integration in the fastai2 library. In the previous version, this was *super* tiresome to do and didn't provide much wiggleroom, especially for image/data specific augmentation. The kernel ([here](https://www.kaggle.com/muellerzr/fastai2-tabular-vision-starter-kernel)) describes the new process. It can be summed up as we make a `DataLoader` of `DataLoaders`. `fastai2` calls the shuffled batches via index's, so we just need to ensure those index's always match up. This can very easily be expanded to not only `n` `DataLoaders` but also can work with *any* `DataLoader` itself, as you're pulling and merging the raw transformed batches *inside* the fastai pipeline. If there are any comments or questions about the process, feel free to ask here!\n\nHere's the idea (for this competition specifically):\n\n## The Pipeline\nHere is an outline of how you go about doing this:\n1. Make your `tab` and `vis` `DataLoaders`\n    (`vis` = Vision, `tab` = Tabular)\n2. Combine them together into a `Hybrid DataLoader`\n3. Adjust your own `test_dl` framework how you choose\n4. Train\n\n## The Code:\nNow let's talk about the code. For our \"DataLoader\", it won't inherit the `DataLoader` class (hence the quotes around it). Instead we'll give it the *minimal similar behavior* to a `DataLoader` that is needed, and have everything else work internally. Specifically, these functionalities:\n* `FakeLoader`\n* `__len__`\n* `__iter__`\n* `one_batch`\n* `show_batch`\n* `shuffle_fn`\n* `to`\n\nNow to build this I'm going to walk us through it with `@patch` from the `fastcore` library. Basically this lets us lazily define the class as we go, so don't get confused to why it's all in more than one block.\n\n## `__init__` and `FakeLoader`\nThe `__init__` for our model needs to store 5 items, the `device` we're running on, our two `DataLoaders` we're passing in, a `count`, a `_FakeLoader`, and our new shuffle function (for now this will be undefined, we'll discuss it more in a moment). Also, `FakeLoader` is used during the `__iter__`, see the regular `DataLoader` source code to see it there:\n\n```python\nfrom fastai2.data.load import _FakeLoader, _loaders\nclass MixedDL():\n    def __init__(self, tab_dl:TabDataLoader, vis_dl:TfmdDL, device='cuda:0'):\n        \"Stores away `tab_dl` and `vis_dl`, and overrides `shuffle_fn`\"\n        self.device = device\n        tab_dl.shuffle_fn = self.shuffle_fn\n        vis_dl.shuffle_fn = self.shuffle_fn\n        self.dls = [tab_dl, vis_dl]\n        self.count = 0\n        self.fake_l = _FakeLoader(self, False, 0, 0)\n```\n\n## `shuffle_fn`\nNow we'll look at the `shuffle_fn` there. What needs to have happen? The `shuffle_fn` returns a list of index's for us to use, that's stored inside of `self.rng`, and we want those index's to change every 2 times we call the `shuffle_fn` (as we call it for each of our internal `DataLoaders`), to ensure that both are mapped out to the same index's for preparing our batch. This is what that looks like:\n```python\n@patch   \ndef shuffle_fn(x:MixedDL, idxs):\n        \"Generates a new `rng` based upon which `DataLoader` is called\"\n        if self.count == 0: # if we haven't generated an rng yet\n            self.rng = self.dls[0].rng.sample(idxs, len(idxs))\n            self.count += 1\n            return self.rng\n        else:\n            self.count = 0\n            return self.rng\n```\nThis is **all** that's needed to ensure that all of our batches get shuffled together. And if you're using more than two, count is just equal to `n` internal `DataLoaders`.\n\nWhile we're at it, we'll take care of two other functions, the `__len__` attribute and the `to` function. `__len__` just needs to grab the length of *one* of our `DataLoaders`, and `to` just returns the name of our device:\n\n```python\n@patch \ndef __len__(x:HybridDL): return len(x.dls[0])\n\n@patch\ndef to(x:HybridDL, device): x.device = device\n```\n\n## `__iter__`\nNow let's move into something a bit more complex, the iterator. Now, our iterator needs to take all of our batches from our loaders and perform the `after_batch` transform for those outputs from *their respective `DataLoader`* before finally being put into a batch, also moving each to the `device`. While this may look scary, the `_loaders` etc is all the same as it is from the `DataLoaders` class, so it's just how we access them:\n\n```python\n@patch\ndef __iter__(dl:MixedDL):\n    \"Iterate over your `DataLoader`\"\n    z = zip(*[_loaders[i.fake_l.num_workers==0](i.fake_l) for i in dl.dls])\n    for b in z:\n        if dl.device is not None: \n            b = to_device(b, dl.device)\n        batch = []\n        batch.extend(dl.dls[0].after_batch(b[0])[:2]) # tabular cat and cont\n        batch.append(dl.dls[1].after_batch(b[1][0])) # Image\n        try: # In case the data is unlabelled\n            batch.append(b[1][1]) # y\n            yield tuple(batch)\n        except:\n            yield tuple(batch)\n```\nNotice the device is adjusted recursively before we move to the batch transforms (this is how `fastai` moves them all to the GPU)\n\n## `one_batch`\n\nAlright, so we can build it, iterate it, now how do we get our good ol' fashion `one_batch`? Quite easily. We call `fake_l.no_multiproc()` (which so you know, that means we temporarily adjust the `num_workers` in our `DataLoader` to zero) and grab the first batch, while also discarding any iterators the `DataLoader` may have (as `first` calls `next(iter(dl))`):\n```python\n@patch\ndef one_batch(x:MixedDL):\n    \"Grab a batch from the `DataLoader`\"\n    with x.fake_l.no_multiproc(): res = first(x)\n    if hasattr(x, 'it'): delattr(x, 'it')\n    return res\n```\nYou *may* or may not get an exception error, this can be safely ignored. Your batch now returns as `[cat, cont, im, y]`\n\n## `show_batch`\n\nNext up is probably the easiest out of all of the functions. All we're wanting to do here is in each `DataLoader`, call `show_batch`. It's as simple as it sounds:\n```python\n@patch\ndef show_batch(x:MixedDL):\n    \"Show a batch from multiple `DataLoaders`\"\n    for dl in x.dls:\n        dl.show_batch()\n```\nHere's an example output:\n\n![](https://storage.googleapis.com/kaggle-forum-message-attachments/883597/15850/26cee5dec222e423748e26b14ae2776011974ae6.jpeg)\n\nAnd that's **all** that's needed to start training and have *all* the functionalities of `fastai` while bringing in the various `DataTypes`. So they key that made this entire thing possible is due to how `fastai` does the `shuffle_fn`, and the fact they are indices. \n\n## `test_dl`\n\nThe last thing I'll show is how to do the `test_dl`. When you're making these ideally you build the entire Image and Tabular `dls`, which gives you access to the `.test_dl` function. From there, simply do something like:\n```python\nim_test = vis_dl.test_dl(test_df)\ntab_test = tab_dl.test_dl(test_df)\ntest_dl = MixedDL(tab_test, im_test)\n```\nAnd you're good to go! The main reason we don't have to worry about enabling shuffling, etc is due to the fact it's done on the *interior* `DataLoader` level.\n\nLet me know any questions you have!"
  }
}