{
  "id": 156163,
  "title": "Some findings from analysing DCT table changes during embedding",
  "url": "/competitions/alaska2-image-steganalysis/discussion/156163",
  "author_name": "",
  "post_date": "2020-06-04T16:28:15.152734900Z",
  "votes": 12,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I've just published this <a href=\"https://www.kaggle.com/vzaguskin/dct-changes-eda-for-alaska/\">EDA kernel</a> where I compare DCT values of each original cover image with its embeddings. Here are some key findings:</p>\n\n<ol>\n<li>The payloads can vary quite significantly in size</li>\n<li>JMiPOD pictures have generally mode bits altered</li>\n<li>Higher quality images have mode bits altered</li>\n<li>For some images there are no changes at all.</li>\n</ol>\n\n<p>For 4, there are totally 250 images that are absolutely identical to cover, but marked as having some embeddings. And also some number of images where just a few coefficients changed, so only few bits of information can maximum be embedded.</p>\n\n<p>I think that is theoretically possible for a very small payload to not change anything in the image. But that obviously makes it totally impossible to detect.</p>",
  "messages": [
    {
      "id": "874096",
      "postDate": "06/04/2020 16:28:15",
      "content": "<p>I've just published this <a href=\"https://www.kaggle.com/vzaguskin/dct-changes-eda-for-alaska/\">EDA kernel</a> where I compare DCT values of each original cover image with its embeddings. Here are some key findings:</p>\n\n<ol>\n<li>The payloads can vary quite significantly in size</li>\n<li>JMiPOD pictures have generally mode bits altered</li>\n<li>Higher quality images have mode bits altered</li>\n<li>For some images there are no changes at all.</li>\n</ol>\n\n<p>For 4, there are totally 250 images that are absolutely identical to cover, but marked as having some embeddings. And also some number of images where just a few coefficients changed, so only few bits of information can maximum be embedded.</p>\n\n<p>I think that is theoretically possible for a very small payload to not change anything in the image. But that obviously makes it totally impossible to detect.</p>",
      "rawMarkdown": "I've just published this [EDA kernel](https://www.kaggle.com/vzaguskin/dct-changes-eda-for-alaska/) where I compare DCT values of each original cover image with its embeddings. Here are some key findings:\n\n1. The payloads can vary quite significantly in size\n2. JMiPOD pictures have generally mode bits altered\n3. Higher quality images have mode bits altered\n4. For some images there are no changes at all.\n\nFor 4, there are totally 250 images that are absolutely identical to cover, but marked as having some embeddings. And also some number of images where just a few coefficients changed, so only few bits of information can maximum be embedded.\n\nI think that is theoretically possible for a very small payload to not change anything in the image. But that obviously makes it totally impossible to detect.",
      "votes": null
    },
    {
      "id": "874269",
      "postDate": "06/04/2020 19:08:59",
      "content": "<p>Hello there,</p>\n\n<p>That is a good remark / findings.\nIn fact as stated in the <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/data\">data description page</a> : \n<code>The payload (message length) is adjusted such that the \"difficulty\" is approximately the same regardless the content of the image. Images with smooth content are used to hide shorter messages while highly textured images will be used to hide more secret bits. The payload is adjusted in the same manner for testing and training sets.</code>\n<code>The average message length is 0.4 bit per non-zero AC DCT coefficient.</code>\nThere are two main reason why higher quality images have subjected to more changes. First, the number of non-zero DCT coefficients is higher. Since we measure we the payload generally in number of bit embedded per non zero DCT coefficients (mostly for old historical reason).\nSecond, and most important, we adjust the message length to have a \"detection difficulty\" that is more or less the same for each image we tend to embed more data into:\n(1) image compressed with higher JPEG quality (that preserve more details and, hence, room for data hiding) \n(2) image with more \"content\" and that contains area more \"complex\" have more room to hide data that badly blurred images</p>",
      "rawMarkdown": "Hello there,\n\nThat is a good remark / findings.\nIn fact as stated in the [data description page](https://www.kaggle.com/c/alaska2-image-steganalysis/data) : \n`The payload (message length) is adjusted such that the \"difficulty\" is approximately the same regardless the content of the image. Images with smooth content are used to hide shorter messages while highly textured images will be used to hide more secret bits. The payload is adjusted in the same manner for testing and training sets.`\n`The average message length is 0.4 bit per non-zero AC DCT coefficient.`\nThere are two main reason why higher quality images have subjected to more changes. First, the number of non-zero DCT coefficients is higher. Since we measure we the payload generally in number of bit embedded per non zero DCT coefficients (mostly for old historical reason).\nSecond, and most important, we adjust the message length to have a \"detection difficulty\" that is more or less the same for each image we tend to embed more data into:\n(1) image compressed with higher JPEG quality (that preserve more details and, hence, room for data hiding) \n(2) image with more \"content\" and that contains area more \"complex\" have more room to hide data that badly blurred images",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 874269,
      "author_name": "remicogranne",
      "author_url": "",
      "post_date": "06/04/2020 19:08:59",
      "content": "<p>Hello there,</p>\n\n<p>That is a good remark / findings.\nIn fact as stated in the <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/data\">data description page</a> : \n<code>The payload (message length) is adjusted such that the \"difficulty\" is approximately the same regardless the content of the image. Images with smooth content are used to hide shorter messages while highly textured images will be used to hide more secret bits. The payload is adjusted in the same manner for testing and training sets.</code>\n<code>The average message length is 0.4 bit per non-zero AC DCT coefficient.</code>\nThere are two main reason why higher quality images have subjected to more changes. First, the number of non-zero DCT coefficients is higher. Since we measure we the payload generally in number of bit embedded per non zero DCT coefficients (mostly for old historical reason).\nSecond, and most important, we adjust the message length to have a \"detection difficulty\" that is more or less the same for each image we tend to embed more data into:\n(1) image compressed with higher JPEG quality (that preserve more details and, hence, room for data hiding) \n(2) image with more \"content\" and that contains area more \"complex\" have more room to hide data that badly blurred images</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "874096": "I've just published this [EDA kernel](https://www.kaggle.com/vzaguskin/dct-changes-eda-for-alaska/) where I compare DCT values of each original cover image with its embeddings. Here are some key findings:\n\n1. The payloads can vary quite significantly in size\n2. JMiPOD pictures have generally mode bits altered\n3. Higher quality images have mode bits altered\n4. For some images there are no changes at all.\n\nFor 4, there are totally 250 images that are absolutely identical to cover, but marked as having some embeddings. And also some number of images where just a few coefficients changed, so only few bits of information can maximum be embedded.\n\nI think that is theoretically possible for a very small payload to not change anything in the image. But that obviously makes it totally impossible to detect.",
    "874269": "Hello there,\n\nThat is a good remark / findings.\nIn fact as stated in the [data description page](https://www.kaggle.com/c/alaska2-image-steganalysis/data) : \n`The payload (message length) is adjusted such that the \"difficulty\" is approximately the same regardless the content of the image. Images with smooth content are used to hide shorter messages while highly textured images will be used to hide more secret bits. The payload is adjusted in the same manner for testing and training sets.`\n`The average message length is 0.4 bit per non-zero AC DCT coefficient.`\nThere are two main reason why higher quality images have subjected to more changes. First, the number of non-zero DCT coefficients is higher. Since we measure we the payload generally in number of bit embedded per non zero DCT coefficients (mostly for old historical reason).\nSecond, and most important, we adjust the message length to have a \"detection difficulty\" that is more or less the same for each image we tend to embed more data into:\n(1) image compressed with higher JPEG quality (that preserve more details and, hence, room for data hiding) \n(2) image with more \"content\" and that contains area more \"complex\" have more room to hide data that badly blurred images"
  },
  "source": "meta"
}