{
  "id": 103475,
  "title": "problem of validation dataset in fastai",
  "url": "/competitions/aptos2019-blindness-detection/discussion/103475",
  "author_name": "",
  "post_date": "2019-08-09T13:37:37.633598200Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>In fastai, the validation dataset is acquired randomly as code below, does anyone kown how to get validation dataset with a certain path ?</p>\n\n<p>data = (\n    train.split_by_rand_pct(0.2)\n    .label_from_df(cols='diagnosis', label_cls=FloatList)\n    .add_test(test)\n    .transform(tfms, size=SIZE)\n    .databunch(path=Path('.'), bs=BS).normalize(imagenet_stats)\n)</p>",
  "messages": [
    {
      "id": "595629",
      "postDate": "08/09/2019 13:37:37",
      "content": "<p>In fastai, the validation dataset is acquired randomly as code below, does anyone kown how to get validation dataset with a certain path ?</p>\n\n<p>data = (\n    train.split_by_rand_pct(0.2)\n    .label_from_df(cols='diagnosis', label_cls=FloatList)\n    .add_test(test)\n    .transform(tfms, size=SIZE)\n    .databunch(path=Path('.'), bs=BS).normalize(imagenet_stats)\n)</p>",
      "rawMarkdown": "In fastai, the validation dataset is acquired randomly as code below, does anyone kown how to get validation dataset with a certain path ?\n\ndata = (\n    train.split_by_rand_pct(0.2)\n    .label_from_df(cols='diagnosis', label_cls=FloatList)\n    .add_test(test)\n    .transform(tfms, size=SIZE)\n    .databunch(path=Path('.'), bs=BS).normalize(imagenet_stats)\n)",
      "votes": null
    },
    {
      "id": "595634",
      "postDate": "08/09/2019 13:43:20",
      "content": "<p>You can add an is_valid column in your dataframe and use split_from_df (more flexible imo): <a href=\"https://docs.fast.ai/data_block.html#ItemList.split_from_df\">https://docs.fast.ai/data_block.html#ItemList.split_from_df</a></p>\n\n<p>Or you can use split_by_folder if your validation set is in a separate folder.</p>",
      "rawMarkdown": "You can add an is_valid column in your dataframe and use split_from_df (more flexible imo): https://docs.fast.ai/data_block.html#ItemList.split_from_df\n\nOr you can use split_by_folder if your validation set is in a separate folder.",
      "votes": null
    },
    {
      "id": "595657",
      "postDate": "08/09/2019 14:05:28",
      "content": "<p>thank you !</p>",
      "rawMarkdown": "thank you !",
      "votes": null
    },
    {
      "id": "794363",
      "postDate": "04/01/2020 18:59:51",
      "content": "<p>As SD said, an easy way is to use .split_from_df</p>\n\n<p>You create an extra column in the df as indicated here <a href=\"https://docs.fast.ai/data_block.html#ItemList.split_from_df\">https://docs.fast.ai/data_block.html#ItemList.split_from_df</a></p>\n\n<p>How you decide what goes into Valid is up to you.</p>\n\n<p>Here is a way to integrate it.</p>\n\n<p>```\ndata = (ImageList.from_df(df=train_mini,cols=\"file_name\",path=pImagesDir) </p>\n\n<h1>Where to find the data? -&gt; in path and its subfolders</h1>\n\n<pre><code>    .split_from_df(col=\"is_valid\")              \n</code></pre>\n\n<h1>How to split in train/valid? -&gt; use the folders</h1>\n\n<pre><code>    .label_from_df(cols='category_id')            \n</code></pre>\n\n<h1>How to label? -&gt; depending on the folder of the filenames</h1>\n\n<pre><code>    #.add_test_folder()              \n</code></pre>\n\n<h1>Optionally add a test set (here default name is test)</h1>\n\n<pre><code>    .transform(tfms, size=images_size)       \n</code></pre>\n\n<h1>Data augmentation? -&gt; use tfms with a size of 64</h1>\n\n<pre><code>    .databunch())                   \n</code></pre>\n\n<h1>Finally? -&gt; use the defaults for conversion to ImageDataBunch</h1>\n\n<p>data=data.normalize(imagenet_stats)</p>\n\n<p>```</p>",
      "rawMarkdown": "As SD said, an easy way is to use .split_from_df\n\nYou create an extra column in the df as indicated here https://docs.fast.ai/data_block.html#ItemList.split_from_df\n\nHow you decide what goes into Valid is up to you.\n\nHere is a way to integrate it.\n\n```\ndata = (ImageList.from_df(df=train_mini,cols=\"file_name\",path=pImagesDir) \n#Where to find the data? -&gt; in path and its subfolders\n        .split_from_df(col=\"is_valid\")              \n#How to split in train/valid? -&gt; use the folders\n        .label_from_df(cols='category_id')            \n#How to label? -&gt; depending on the folder of the filenames\n        #.add_test_folder()              \n#Optionally add a test set (here default name is test)\n        .transform(tfms, size=images_size)       \n#Data augmentation? -&gt; use tfms with a size of 64\n        .databunch())                   \n#Finally? -&gt; use the defaults for conversion to ImageDataBunch\ndata=data.normalize(imagenet_stats)\n\n\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 595634,
      "author_name": "sdoria",
      "author_url": "",
      "post_date": "08/09/2019 13:43:20",
      "content": "<p>You can add an is_valid column in your dataframe and use split_from_df (more flexible imo): <a href=\"https://docs.fast.ai/data_block.html#ItemList.split_from_df\">https://docs.fast.ai/data_block.html#ItemList.split_from_df</a></p>\n\n<p>Or you can use split_by_folder if your validation set is in a separate folder.</p>",
      "votes": null,
      "replies": [
        {
          "id": 595657,
          "author_name": "dldmw579",
          "author_url": "",
          "post_date": "08/09/2019 14:05:28",
          "content": "<p>thank you !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 794363,
      "author_name": "mrruben",
      "author_url": "",
      "post_date": "04/01/2020 18:59:51",
      "content": "<p>As SD said, an easy way is to use .split_from_df</p>\n\n<p>You create an extra column in the df as indicated here <a href=\"https://docs.fast.ai/data_block.html#ItemList.split_from_df\">https://docs.fast.ai/data_block.html#ItemList.split_from_df</a></p>\n\n<p>How you decide what goes into Valid is up to you.</p>\n\n<p>Here is a way to integrate it.</p>\n\n<p>```\ndata = (ImageList.from_df(df=train_mini,cols=\"file_name\",path=pImagesDir) </p>\n\n<h1>Where to find the data? -&gt; in path and its subfolders</h1>\n\n<pre><code>    .split_from_df(col=\"is_valid\")              \n</code></pre>\n\n<h1>How to split in train/valid? -&gt; use the folders</h1>\n\n<pre><code>    .label_from_df(cols='category_id')            \n</code></pre>\n\n<h1>How to label? -&gt; depending on the folder of the filenames</h1>\n\n<pre><code>    #.add_test_folder()              \n</code></pre>\n\n<h1>Optionally add a test set (here default name is test)</h1>\n\n<pre><code>    .transform(tfms, size=images_size)       \n</code></pre>\n\n<h1>Data augmentation? -&gt; use tfms with a size of 64</h1>\n\n<pre><code>    .databunch())                   \n</code></pre>\n\n<h1>Finally? -&gt; use the defaults for conversion to ImageDataBunch</h1>\n\n<p>data=data.normalize(imagenet_stats)</p>\n\n<p>```</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "595629": "In fastai, the validation dataset is acquired randomly as code below, does anyone kown how to get validation dataset with a certain path ?\n\ndata = (\n    train.split_by_rand_pct(0.2)\n    .label_from_df(cols='diagnosis', label_cls=FloatList)\n    .add_test(test)\n    .transform(tfms, size=SIZE)\n    .databunch(path=Path('.'), bs=BS).normalize(imagenet_stats)\n)",
    "595634": "You can add an is_valid column in your dataframe and use split_from_df (more flexible imo): https://docs.fast.ai/data_block.html#ItemList.split_from_df\n\nOr you can use split_by_folder if your validation set is in a separate folder.",
    "595657": "thank you !",
    "794363": "As SD said, an easy way is to use .split_from_df\n\nYou create an extra column in the df as indicated here https://docs.fast.ai/data_block.html#ItemList.split_from_df\n\nHow you decide what goes into Valid is up to you.\n\nHere is a way to integrate it.\n\n```\ndata = (ImageList.from_df(df=train_mini,cols=\"file_name\",path=pImagesDir) \n#Where to find the data? -&gt; in path and its subfolders\n        .split_from_df(col=\"is_valid\")              \n#How to split in train/valid? -&gt; use the folders\n        .label_from_df(cols='category_id')            \n#How to label? -&gt; depending on the folder of the filenames\n        #.add_test_folder()              \n#Optionally add a test set (here default name is test)\n        .transform(tfms, size=images_size)       \n#Data augmentation? -&gt; use tfms with a size of 64\n        .databunch())                   \n#Finally? -&gt; use the defaults for conversion to ImageDataBunch\ndata=data.normalize(imagenet_stats)\n\n\n```"
  },
  "source": "meta"
}