{
  "id": 146699,
  "title": "If you want more training data ",
  "url": "/competitions/alaska2-image-steganalysis/discussion/146699",
  "author_name": "",
  "post_date": "2020-04-28T06:54:52.173106200Z",
  "votes": 8,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I also forgot to mention (I may add this information in data part) ; but it worth mentionning.\nAll the image from the training set (of course not the testing set) have been obtained with images from ALASKA image dataset (<a href=\"https://alaska.utt.fr\">https://alaska.utt.fr</a> ; raw images are also provided as well as the development script).\nFeel free to download those image and do the embedding on your own if you feel like the 320,000 training set is not big enough.</p>",
  "messages": [
    {
      "id": "824117",
      "postDate": "04/28/2020 06:54:52",
      "content": "<p>I also forgot to mention (I may add this information in data part) ; but it worth mentionning.\nAll the image from the training set (of course not the testing set) have been obtained with images from ALASKA image dataset (<a href=\"https://alaska.utt.fr\">https://alaska.utt.fr</a> ; raw images are also provided as well as the development script).\nFeel free to download those image and do the embedding on your own if you feel like the 320,000 training set is not big enough.</p>",
      "rawMarkdown": "I also forgot to mention (I may add this information in data part) ; but it worth mentionning.\nAll the image from the training set (of course not the testing set) have been obtained with images from ALASKA image dataset (https://alaska.utt.fr ; raw images are also provided as well as the development script).\nFeel free to download those image and do the embedding on your own if you feel like the 320,000 training set is not big enough.",
      "votes": null
    },
    {
      "id": "824466",
      "postDate": "04/28/2020 11:33:02",
      "content": "<p>Hi <a href=\"/remicogranne\">@remicogranne</a> I couldn't find such dataset (<a href=\"https://alaska.utt.fr/#material\">https://alaska.utt.fr/#material</a> is linked to this page) could you please point where are the raw images? thanks!</p>",
      "rawMarkdown": "Hi @remicogranne I couldn't find such dataset (https://alaska.utt.fr/#material is linked to this page) could you please point where are the raw images? thanks!",
      "votes": null
    },
    {
      "id": "824521",
      "postDate": "04/28/2020 12:26:34",
      "content": "<p>You are right and I did a mistake by posted this discussion.\nIn fact we wanted to provide those addition dataset. Yet we believe it may be more difficult for all competitors to access RAW datasets, developement and embedding scripts. I will discuss this with my colleagues and kaggle administators before deciding.\nsorry for the confusion</p>",
      "rawMarkdown": "You are right and I did a mistake by posted this discussion.\nIn fact we wanted to provide those addition dataset. Yet we believe it may be more difficult for all competitors to access RAW datasets, developement and embedding scripts. I will discuss this with my colleagues and kaggle administators before deciding.\nsorry for the confusion",
      "votes": null
    },
    {
      "id": "824560",
      "postDate": "04/28/2020 13:01:13",
      "content": "<p><a href=\"/remicogranne\">@remicogranne</a>, maybe TIFF files can be easier to access?</p>",
      "rawMarkdown": "remicogranne, maybe TIFF files can be easier to access?",
      "votes": null
    },
    {
      "id": "825064",
      "postDate": "04/28/2020 18:57:20",
      "content": "<p>Hello there,</p>\n\n<p>I apologize for the confusion. We have discussed with my colleagues and with kaggle administrators. The goal is to prevent the competition shifting from machine learning challenge to data collection contest. \nWe believe that with 225.000 stego image and 75.000 cover the dataset is quite large enough. </p>\n\n<p>If at some point it seems that the dataset is too narrow and expert from data hiding field with the ability to generate more training example have an advantages we will provide more dataset.</p>\n\n<p>Again, I apologize for the confusion and will delete this discussion in a few days to prevent creating more troubles 😿 </p>",
      "rawMarkdown": "Hello there,\n\nI apologize for the confusion. We have discussed with my colleagues and with kaggle administrators. The goal is to prevent the competition shifting from machine learning challenge to data collection contest. \nWe believe that with 225.000 stego image and 75.000 cover the dataset is quite large enough. \n\nIf at some point it seems that the dataset is too narrow and expert from data hiding field with the ability to generate more training example have an advantages we will provide more dataset.\n\nAgain, I apologize for the confusion and will delete this discussion in a few days to prevent creating more troubles 😿",
      "votes": null
    },
    {
      "id": "825071",
      "postDate": "04/28/2020 19:01:09",
      "content": "<p>wise decision! thanks for the quick response</p>",
      "rawMarkdown": "wise decision! thanks for the quick response",
      "votes": null
    },
    {
      "id": "843713",
      "postDate": "05/12/2020 08:07:29",
      "content": "<p><a href=\"/remicogranne\">@remicogranne</a> </p>\n\n<p>Can you clarify your clarification? The rules currently say external data is allowed. If it is allowed, many will use it, overcoming any difficulties in collection and processing. A goal to prevent the competition shifting from a machine learning challenge to data collection contest is a good goal. I cease participating in challenges where others have the resources to train on 50x data.</p>\n\n<p>In this thread, are you a) saying that external datasets such as Alaska1 and IStego100k are are not allowed, or b) merely that they are allowed but the sponsors and organisers are not formally providing such data.</p>\n\n<p>The hypothetical \"If at some point it seems that the dataset is too narrow and expert from data hiding field with the ability to generate more training example have an advantages we will provide more dataset\" seems probable. Here, training on 90% of data scores better than training on 80% of data, so I imagine training on 110% would perform better again.</p>",
      "rawMarkdown": "remicogranne \n\nCan you clarify your clarification? The rules currently say external data is allowed. If it is allowed, many will use it, overcoming any difficulties in collection and processing. A goal to prevent the competition shifting from a machine learning challenge to data collection contest is a good goal. I cease participating in challenges where others have the resources to train on 50x data.\n\nIn this thread, are you a) saying that external datasets such as Alaska1 and IStego100k are are not allowed, or b) merely that they are allowed but the sponsors and organisers are not formally providing such data.\n\nThe hypothetical \"If at some point it seems that the dataset is too narrow and expert from data hiding field with the ability to generate more training example have an advantages we will provide more dataset\" seems probable. Here, training on 90% of data scores better than training on 80% of data, so I imagine training on 110% would perform better again.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 824466,
      "author_name": "jesucristo",
      "author_url": "",
      "post_date": "04/28/2020 11:33:02",
      "content": "<p>Hi <a href=\"/remicogranne\">@remicogranne</a> I couldn't find such dataset (<a href=\"https://alaska.utt.fr/#material\">https://alaska.utt.fr/#material</a> is linked to this page) could you please point where are the raw images? thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 824521,
      "author_name": "remicogranne",
      "author_url": "",
      "post_date": "04/28/2020 12:26:34",
      "content": "<p>You are right and I did a mistake by posted this discussion.\nIn fact we wanted to provide those addition dataset. Yet we believe it may be more difficult for all competitors to access RAW datasets, developement and embedding scripts. I will discuss this with my colleagues and kaggle administators before deciding.\nsorry for the confusion</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 824560,
      "author_name": "yousfi",
      "author_url": "",
      "post_date": "04/28/2020 13:01:13",
      "content": "<p><a href=\"/remicogranne\">@remicogranne</a>, maybe TIFF files can be easier to access?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 825064,
      "author_name": "remicogranne",
      "author_url": "",
      "post_date": "04/28/2020 18:57:20",
      "content": "<p>Hello there,</p>\n\n<p>I apologize for the confusion. We have discussed with my colleagues and with kaggle administrators. The goal is to prevent the competition shifting from machine learning challenge to data collection contest. \nWe believe that with 225.000 stego image and 75.000 cover the dataset is quite large enough. </p>\n\n<p>If at some point it seems that the dataset is too narrow and expert from data hiding field with the ability to generate more training example have an advantages we will provide more dataset.</p>\n\n<p>Again, I apologize for the confusion and will delete this discussion in a few days to prevent creating more troubles 😿 </p>",
      "votes": null,
      "replies": [
        {
          "id": 825071,
          "author_name": "jesucristo",
          "author_url": "",
          "post_date": "04/28/2020 19:01:09",
          "content": "<p>wise decision! thanks for the quick response</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 843713,
          "author_name": "robga",
          "author_url": "",
          "post_date": "05/12/2020 08:07:29",
          "content": "<p><a href=\"/remicogranne\">@remicogranne</a> </p>\n\n<p>Can you clarify your clarification? The rules currently say external data is allowed. If it is allowed, many will use it, overcoming any difficulties in collection and processing. A goal to prevent the competition shifting from a machine learning challenge to data collection contest is a good goal. I cease participating in challenges where others have the resources to train on 50x data.</p>\n\n<p>In this thread, are you a) saying that external datasets such as Alaska1 and IStego100k are are not allowed, or b) merely that they are allowed but the sponsors and organisers are not formally providing such data.</p>\n\n<p>The hypothetical \"If at some point it seems that the dataset is too narrow and expert from data hiding field with the ability to generate more training example have an advantages we will provide more dataset\" seems probable. Here, training on 90% of data scores better than training on 80% of data, so I imagine training on 110% would perform better again.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "824117": "I also forgot to mention (I may add this information in data part) ; but it worth mentionning.\nAll the image from the training set (of course not the testing set) have been obtained with images from ALASKA image dataset (https://alaska.utt.fr ; raw images are also provided as well as the development script).\nFeel free to download those image and do the embedding on your own if you feel like the 320,000 training set is not big enough.",
    "824466": "Hi @remicogranne I couldn't find such dataset (https://alaska.utt.fr/#material is linked to this page) could you please point where are the raw images? thanks!",
    "824521": "You are right and I did a mistake by posted this discussion.\nIn fact we wanted to provide those addition dataset. Yet we believe it may be more difficult for all competitors to access RAW datasets, developement and embedding scripts. I will discuss this with my colleagues and kaggle administators before deciding.\nsorry for the confusion",
    "824560": "remicogranne, maybe TIFF files can be easier to access?",
    "825064": "Hello there,\n\nI apologize for the confusion. We have discussed with my colleagues and with kaggle administrators. The goal is to prevent the competition shifting from machine learning challenge to data collection contest. \nWe believe that with 225.000 stego image and 75.000 cover the dataset is quite large enough. \n\nIf at some point it seems that the dataset is too narrow and expert from data hiding field with the ability to generate more training example have an advantages we will provide more dataset.\n\nAgain, I apologize for the confusion and will delete this discussion in a few days to prevent creating more troubles 😿",
    "825071": "wise decision! thanks for the quick response",
    "843713": "remicogranne \n\nCan you clarify your clarification? The rules currently say external data is allowed. If it is allowed, many will use it, overcoming any difficulties in collection and processing. A goal to prevent the competition shifting from a machine learning challenge to data collection contest is a good goal. I cease participating in challenges where others have the resources to train on 50x data.\n\nIn this thread, are you a) saying that external datasets such as Alaska1 and IStego100k are are not allowed, or b) merely that they are allowed but the sponsors and organisers are not formally providing such data.\n\nThe hypothetical \"If at some point it seems that the dataset is too narrow and expert from data hiding field with the ability to generate more training example have an advantages we will provide more dataset\" seems probable. Here, training on 90% of data scores better than training on 80% of data, so I imagine training on 110% would perform better again."
  },
  "source": "meta"
}