{
  "id": 313879,
  "title": "Doubling the training set size ",
  "url": "/competitions/ultra-mnist/discussion/313879",
  "author_name": "Ismaila SECK",
  "post_date": "2022-03-19T12:34:35.887000",
  "votes": 9,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Seeing how complex the  task is given the dataset, having more data could benefit our model and help achieve higher accuracy. An easy way to double the size of the training set is to invert the color of each image making recoloring the black parts in white and the white parts in black as illustrated in the following notebook:<br>\n<a href=\"https://www.kaggle.com/ismailaseck/few-lines-to-double-the-training-set-size\" target=\"_blank\">https://www.kaggle.com/ismailaseck/few-lines-to-double-the-training-set-size</a></p>",
  "messages": [
    {
      "id": 1728909,
      "postDate": "2022-03-19T12:34:35.887Z",
      "content": "<p>Seeing how complex the  task is given the dataset, having more data could benefit our model and help achieve higher accuracy. An easy way to double the size of the training set is to invert the color of each image making recoloring the black parts in white and the white parts in black as illustrated in the following notebook:<br>\n<a href=\"https://www.kaggle.com/ismailaseck/few-lines-to-double-the-training-set-size\" target=\"_blank\">https://www.kaggle.com/ismailaseck/few-lines-to-double-the-training-set-size</a></p>",
      "rawMarkdown": "Seeing how complex the  task is given the dataset, having more data could benefit our model and help achieve higher accuracy. An easy way to double the size of the training set is to invert the color of each image making recoloring the black parts in white and the white parts in black as illustrated in the following notebook:\nhttps://www.kaggle.com/ismailaseck/few-lines-to-double-the-training-set-size",
      "votes": 9
    },
    {
      "id": 1731118,
      "postDate": "2022-03-22T02:11:52.217Z",
      "content": "<p>Innovative idea. But robust solution must be generating custom dataset </p>",
      "rawMarkdown": "Innovative idea. But robust solution must be generating custom dataset ",
      "votes": 1,
      "replies": [
        {
          "id": 1731298,
          "postDate": "2022-03-22T08:02:51.443Z",
          "content": "<p>Totally agree. One can even go further by applying on-the-fly modifications to a custom dataset.</p>",
          "rawMarkdown": "Totally agree. One can even go further by applying on-the-fly modifications to a custom dataset."
        },
        {
          "id": 1731471,
          "postDate": "2022-03-22T12:00:43.847Z",
          "content": "<p>But we can't generate a custom dataset right ? I mean it's against the rules ?</p>",
          "rawMarkdown": "But we can't generate a custom dataset right ? I mean it's against the rules ?"
        },
        {
          "id": 1731518,
          "postDate": "2022-03-22T12:48:08.793Z",
          "content": "<p>From what I understood, as far as the dataset you used to create a custom dataset is available at no cost, it is allowed. Maybe I misunderstood the rules.</p>",
          "rawMarkdown": "From what I understood, as far as the dataset you used to create a custom dataset is available at no cost, it is allowed. Maybe I misunderstood the rules."
        },
        {
          "id": 1741463,
          "postDate": "2022-03-31T20:09:44.883Z",
          "content": "<p>Data augmentation is a standart part of DL approach.<br>\nI'm thinking if you did not carefully craft your custom modifications and keep it at randomly applicable transformations of data, it should be valid.</p>\n<p>Adding image negatives is also such an option of data augmentation.</p>\n<p>Is this line of thinking correct? <a href=\"https://www.kaggle.com/dkgupta90\" target=\"_blank\">@dkgupta90</a> <a href=\"https://www.kaggle.com/abhishek\" target=\"_blank\">@abhishek</a> </p>",
          "rawMarkdown": "Data augmentation is a standart part of DL approach.\nI'm thinking if you did not carefully craft your custom modifications and keep it at randomly applicable transformations of data, it should be valid.\n\nAdding image negatives is also such an option of data augmentation.\n\nIs this line of thinking correct? @dkgupta90 @abhishek "
        }
      ]
    },
    {
      "id": 1729252,
      "postDate": "2022-03-19T20:24:48.490Z",
      "content": "<p>Great idea! I think it will be really useful.</p>",
      "rawMarkdown": "Great idea! I think it will be really useful.",
      "votes": 1,
      "replies": [
        {
          "id": 1731297,
          "postDate": "2022-03-22T08:01:27.570Z",
          "content": "<p>Thank you .</p>",
          "rawMarkdown": "Thank you ."
        }
      ]
    },
    {
      "id": 1734360,
      "postDate": "2022-03-25T08:57:31.793Z",
      "content": "<p>Can someone show the code for it, also if possible how to do it with fastai</p>",
      "rawMarkdown": "Can someone show the code for it, also if possible how to do it with fastai",
      "replies": [
        {
          "id": 1734390,
          "postDate": "2022-03-25T09:48:47.587Z",
          "content": "<p>You can find below the code to invert black and white in images: <a href=\"https://www.kaggle.com/ismailaseck/few-lines-to-double-the-training-set-size\" target=\"_blank\">https://www.kaggle.com/ismailaseck/few-lines-to-double-the-training-set-size</a>. <br>\nIf you </p>",
          "rawMarkdown": "You can find below the code to invert black and white in images: https://www.kaggle.com/ismailaseck/few-lines-to-double-the-training-set-size. \nIf you "
        },
        {
          "id": 1734393,
          "postDate": "2022-03-25T09:51:09.657Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1734396,
          "postDate": "2022-03-25T09:52:51.177Z",
          "content": "<p>got it, thanks</p>",
          "rawMarkdown": "got it, thanks"
        },
        {
          "id": 1734415,
          "postDate": "2022-03-25T10:27:07.630Z",
          "content": "<p>You're welcome 🙂</p>",
          "rawMarkdown": "You're welcome 🙂"
        },
        {
          "id": 1740674,
          "postDate": "2022-03-31T07:14:16.247Z",
          "content": "<p>i have created a dataset with 1000px resolution and uploaded here:</p>\n<p><a href=\"https://www.kaggle.com/datasets/abhishekdasani/ultramnist-1000px-train-doubled\" target=\"_blank\">https://www.kaggle.com/datasets/abhishekdasani/ultramnist-1000px-train-doubled</a></p>",
          "rawMarkdown": "i have created a dataset with 1000px resolution and uploaded here:\n\nhttps://www.kaggle.com/datasets/abhishekdasani/ultramnist-1000px-train-doubled",
          "votes": 1
        },
        {
          "id": 1741234,
          "postDate": "2022-03-31T16:28:56.920Z",
          "content": "<p>Awesome, thank you for sharing!</p>",
          "rawMarkdown": "Awesome, thank you for sharing!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1731118,
      "author_name": "Asapanna Rakesh",
      "author_url": "",
      "post_date": "2022-03-22T02:11:52.217000",
      "content": "<p>Innovative idea. But robust solution must be generating custom dataset </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1731298,
          "author_name": "Ismaila SECK",
          "author_url": "",
          "post_date": "2022-03-22T08:02:51.443000",
          "content": "<p>Totally agree. One can even go further by applying on-the-fly modifications to a custom dataset.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1731471,
          "author_name": "Yohann Lereclus",
          "author_url": "",
          "post_date": "2022-03-22T12:00:43.847000",
          "content": "<p>But we can't generate a custom dataset right ? I mean it's against the rules ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1731518,
          "author_name": "Ismaila SECK",
          "author_url": "",
          "post_date": "2022-03-22T12:48:08.793000",
          "content": "<p>From what I understood, as far as the dataset you used to create a custom dataset is available at no cost, it is allowed. Maybe I misunderstood the rules.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1741463,
          "author_name": "faraday",
          "author_url": "",
          "post_date": "2022-03-31T20:09:44.883000",
          "content": "<p>Data augmentation is a standart part of DL approach.<br>\nI'm thinking if you did not carefully craft your custom modifications and keep it at randomly applicable transformations of data, it should be valid.</p>\n<p>Adding image negatives is also such an option of data augmentation.</p>\n<p>Is this line of thinking correct? <a href=\"https://www.kaggle.com/dkgupta90\" target=\"_blank\">@dkgupta90</a> <a href=\"https://www.kaggle.com/abhishek\" target=\"_blank\">@abhishek</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1729252,
      "author_name": "dmitry_lessy",
      "author_url": "",
      "post_date": "2022-03-19T20:24:48.490000",
      "content": "<p>Great idea! I think it will be really useful.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1731297,
          "author_name": "Ismaila SECK",
          "author_url": "",
          "post_date": "2022-03-22T08:01:27.570000",
          "content": "<p>Thank you .</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1734360,
      "author_name": "abhishek dasani",
      "author_url": "",
      "post_date": "2022-03-25T08:57:31.793000",
      "content": "<p>Can someone show the code for it, also if possible how to do it with fastai</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1734390,
          "author_name": "Ismaila SECK",
          "author_url": "",
          "post_date": "2022-03-25T09:48:47.587000",
          "content": "<p>You can find below the code to invert black and white in images: <a href=\"https://www.kaggle.com/ismailaseck/few-lines-to-double-the-training-set-size\" target=\"_blank\">https://www.kaggle.com/ismailaseck/few-lines-to-double-the-training-set-size</a>. <br>\nIf you </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1734393,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-03-25T09:51:09.657000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1734396,
          "author_name": "abhishek dasani",
          "author_url": "",
          "post_date": "2022-03-25T09:52:51.177000",
          "content": "<p>got it, thanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1734415,
          "author_name": "Ismaila SECK",
          "author_url": "",
          "post_date": "2022-03-25T10:27:07.630000",
          "content": "<p>You're welcome 🙂</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1740674,
          "author_name": "abhishek dasani",
          "author_url": "",
          "post_date": "2022-03-31T07:14:16.247000",
          "content": "<p>i have created a dataset with 1000px resolution and uploaded here:</p>\n<p><a href=\"https://www.kaggle.com/datasets/abhishekdasani/ultramnist-1000px-train-doubled\" target=\"_blank\">https://www.kaggle.com/datasets/abhishekdasani/ultramnist-1000px-train-doubled</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1741234,
          "author_name": "Ismaila SECK",
          "author_url": "",
          "post_date": "2022-03-31T16:28:56.920000",
          "content": "<p>Awesome, thank you for sharing!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1728909": "Seeing how complex the  task is given the dataset, having more data could benefit our model and help achieve higher accuracy. An easy way to double the size of the training set is to invert the color of each image making recoloring the black parts in white and the white parts in black as illustrated in the following notebook:\nhttps://www.kaggle.com/ismailaseck/few-lines-to-double-the-training-set-size",
    "1731118": "Innovative idea. But robust solution must be generating custom dataset ",
    "1729252": "Great idea! I think it will be really useful.",
    "1734360": "Can someone show the code for it, also if possible how to do it with fastai"
  }
}