{
  "id": 147494,
  "title": "Impact of steganography: do not look only at pixel difference",
  "url": "/competitions/alaska2-image-steganalysis/discussion/147494",
  "author_name": "",
  "post_date": "2020-04-30T19:54:58.261654800Z",
  "votes": 43,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello There,</p>\n\n<p>I have seen many of you looking at the difference between images to \"see\" the difference between cover and stego image. \nI just wanted to state loudly that this is a consequence of data hiding. Indeed, the image are compressed in JPEG and you not be able to embed data in pixel ... after compression it would be unreadable.\nTo answer a question from @vovanf98 I made a small notebook (I apologize, I am terrible at programming) available here, <em>for which I deeply thank @yousfi for his help using external library</em>:\n<a href=\"https://www.kaggle.com/remicogranne/inspect-impact-of-steganography-on-dct-coefs\">https://www.kaggle.com/remicogranne/inspect-impact-of-steganography-on-dct-coefs</a>\nwhich aims to show several basic things:\n(1) data are hidden into DCT coefficients. Looking at pixels is a consequence of data hiding not the hidden data themselves.\n(2) Looking at pixels difference can give you the taste that difference (hidden data) are more or less evenly spread in all color. It is not ! Much more data are hidden in the Y channel of JPEG. Yet you can hardly see this in spatial domain.\n(3) Of course the embedding is adaptive in the sense that it embeds mostly around \"complex areas\", more bits in \"textured images\" and more bits in images compressed with higher JPEG quality factors.</p>\n\n<p>(4) Eventually, I would like share a few question that remain open in the community:\n- Should we do the detection in pixels or in DCT coefficients ? Pixels are more easy to model while the hidden data you want to detect really is hidden in DCT ....\n- We used three different JPEG quality factors (95/90/75) and three different embedding scheme. Is it better to merge everything or should you try to do the learning for each QF and each embedding (on testing you can always identify QF by looking at Qtable and use a fusion function for the embedding...)\n- Similarly, should you merge all images or try to separate them by content ? Noise level ? \nIn other words the best may be trying to be as universal as possible or as targeted as possible ... </p>",
  "messages": [
    {
      "id": "828180",
      "postDate": "04/30/2020 19:54:58",
      "content": "<p>Hello There,</p>\n\n<p>I have seen many of you looking at the difference between images to \"see\" the difference between cover and stego image. \nI just wanted to state loudly that this is a consequence of data hiding. Indeed, the image are compressed in JPEG and you not be able to embed data in pixel ... after compression it would be unreadable.\nTo answer a question from @vovanf98 I made a small notebook (I apologize, I am terrible at programming) available here, <em>for which I deeply thank @yousfi for his help using external library</em>:\n<a href=\"https://www.kaggle.com/remicogranne/inspect-impact-of-steganography-on-dct-coefs\">https://www.kaggle.com/remicogranne/inspect-impact-of-steganography-on-dct-coefs</a>\nwhich aims to show several basic things:\n(1) data are hidden into DCT coefficients. Looking at pixels is a consequence of data hiding not the hidden data themselves.\n(2) Looking at pixels difference can give you the taste that difference (hidden data) are more or less evenly spread in all color. It is not ! Much more data are hidden in the Y channel of JPEG. Yet you can hardly see this in spatial domain.\n(3) Of course the embedding is adaptive in the sense that it embeds mostly around \"complex areas\", more bits in \"textured images\" and more bits in images compressed with higher JPEG quality factors.</p>\n\n<p>(4) Eventually, I would like share a few question that remain open in the community:\n- Should we do the detection in pixels or in DCT coefficients ? Pixels are more easy to model while the hidden data you want to detect really is hidden in DCT ....\n- We used three different JPEG quality factors (95/90/75) and three different embedding scheme. Is it better to merge everything or should you try to do the learning for each QF and each embedding (on testing you can always identify QF by looking at Qtable and use a fusion function for the embedding...)\n- Similarly, should you merge all images or try to separate them by content ? Noise level ? \nIn other words the best may be trying to be as universal as possible or as targeted as possible ... </p>",
      "rawMarkdown": "Hello There,\n\nI have seen many of you looking at the difference between images to \"see\" the difference between cover and stego image. \nI just wanted to state loudly that this is a consequence of data hiding. Indeed, the image are compressed in JPEG and you not be able to embed data in pixel ... after compression it would be unreadable.\nTo answer a question from @vovanf98 I made a small notebook (I apologize, I am terrible at programming) available here, *for which I deeply thank @yousfi for his help using external library*:\n[https://www.kaggle.com/remicogranne/inspect-impact-of-steganography-on-dct-coefs](https://www.kaggle.com/remicogranne/inspect-impact-of-steganography-on-dct-coefs)\nwhich aims to show several basic things:\n(1) data are hidden into DCT coefficients. Looking at pixels is a consequence of data hiding not the hidden data themselves.\n(2) Looking at pixels difference can give you the taste that difference (hidden data) are more or less evenly spread in all color. It is not ! Much more data are hidden in the Y channel of JPEG. Yet you can hardly see this in spatial domain.\n(3) Of course the embedding is adaptive in the sense that it embeds mostly around \"complex areas\", more bits in \"textured images\" and more bits in images compressed with higher JPEG quality factors.\n\n(4) Eventually, I would like share a few question that remain open in the community:\n- Should we do the detection in pixels or in DCT coefficients ? Pixels are more easy to model while the hidden data you want to detect really is hidden in DCT ....\n- We used three different JPEG quality factors (95/90/75) and three different embedding scheme. Is it better to merge everything or should you try to do the learning for each QF and each embedding (on testing you can always identify QF by looking at Qtable and use a fusion function for the embedding...)\n- Similarly, should you merge all images or try to separate them by content ? Noise level ? \nIn other words the best may be trying to be as universal as possible or as targeted as possible ...",
      "votes": null
    },
    {
      "id": "833318",
      "postDate": "05/04/2020 18:12:43",
      "content": "<p>The answer to the question whether to look in the DCT coefficients or in the pixels for a pattern seems to be not so straightforward to answer. Many papers claim that it is common sense to use the same domain that is used for embedding. However, for example in this thesis (<a href=\"http://dde.binghamton.edu/vholub/pdf/Holub_PhD_Dissertation_2014.pdf\">http://dde.binghamton.edu/vholub/pdf/Holub_PhD_Dissertation_2014.pdf</a>) you can find a different conclusion (last chapter). It is claimed that also for DCT embedding it it is easier to detect the embedding in the spatial domain, since most algorithms use a cost function defined in the spatial domain. </p>",
      "rawMarkdown": "The answer to the question whether to look in the DCT coefficients or in the pixels for a pattern seems to be not so straightforward to answer. Many papers claim that it is common sense to use the same domain that is used for embedding. However, for example in this thesis (http://dde.binghamton.edu/vholub/pdf/Holub_PhD_Dissertation_2014.pdf) you can find a different conclusion (last chapter). It is claimed that also for DCT embedding it it is easier to detect the embedding in the spatial domain, since most algorithms use a cost function defined in the spatial domain.",
      "votes": null
    },
    {
      "id": "835301",
      "postDate": "05/06/2020 07:12:37",
      "content": "<p>Hello there,\nI agree with your post, different source will provide different argument whether one should use the domain use for embedding (DCT values) or the one in which it is easier to model the image (spatial / pixels). \nOne thing that is sure is that until now, most steganalysis methods have been using pixels values, IMHO because DCT coefficients are more difficult to interpret / model / analysis. \nThis is true for \"classical\" steganalysis that uses features extraction + supervised learning.\nThis is even more for Deep Learning based staganalysis: applying local convolution over DCT values seems very difficult becase DCT values next to each other are very different.... (hence the proposal in <a href=\"http://dde.binghamton.edu/vholub/pdf/Holub_PhD_Dissertation_2014.pdf\">chapitre 8 of the thesis you cited</a> to use of filters of the DCT without subsampling ....).</p>\n\n<p>I am also very surprised that even without separating image compressed with different quality factors (QF) pretty good results can be obtained ... or perhaps the top scorer are separating those image and train three different networks .... </p>",
      "rawMarkdown": "Hello there,\nI agree with your post, different source will provide different argument whether one should use the domain use for embedding (DCT values) or the one in which it is easier to model the image (spatial / pixels). \nOne thing that is sure is that until now, most steganalysis methods have been using pixels values, IMHO because DCT coefficients are more difficult to interpret / model / analysis. \nThis is true for \"classical\" steganalysis that uses features extraction + supervised learning.\nThis is even more for Deep Learning based staganalysis: applying local convolution over DCT values seems very difficult becase DCT values next to each other are very different.... (hence the proposal in [chapitre 8 of the thesis you cited](http://dde.binghamton.edu/vholub/pdf/Holub_PhD_Dissertation_2014.pdf) to use of filters of the DCT without subsampling ....).\n\nI am also very surprised that even without separating image compressed with different quality factors (QF) pretty good results can be obtained ... or perhaps the top scorer are separating those image and train three different networks ....",
      "votes": null
    },
    {
      "id": "835462",
      "postDate": "05/06/2020 09:28:28",
      "content": "<p>It is possible a network fed with pixels is learning an approximation of the DCT coefficients, and mixed QFs may provide regularisation. I am yet to try it, but stride 8 convolutions in the first layer of an architecture may provide some hinting to counter the fact that DCT values next to each other are very different.</p>",
      "rawMarkdown": "It is possible a network fed with pixels is learning an approximation of the DCT coefficients, and mixed QFs may provide regularisation. I am yet to try it, but stride 8 convolutions in the first layer of an architecture may provide some hinting to counter the fact that DCT values next to each other are very different.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 833318,
      "author_name": "manuelkraus89",
      "author_url": "",
      "post_date": "05/04/2020 18:12:43",
      "content": "<p>The answer to the question whether to look in the DCT coefficients or in the pixels for a pattern seems to be not so straightforward to answer. Many papers claim that it is common sense to use the same domain that is used for embedding. However, for example in this thesis (<a href=\"http://dde.binghamton.edu/vholub/pdf/Holub_PhD_Dissertation_2014.pdf\">http://dde.binghamton.edu/vholub/pdf/Holub_PhD_Dissertation_2014.pdf</a>) you can find a different conclusion (last chapter). It is claimed that also for DCT embedding it it is easier to detect the embedding in the spatial domain, since most algorithms use a cost function defined in the spatial domain. </p>",
      "votes": null,
      "replies": [
        {
          "id": 835301,
          "author_name": "remicogranne",
          "author_url": "",
          "post_date": "05/06/2020 07:12:37",
          "content": "<p>Hello there,\nI agree with your post, different source will provide different argument whether one should use the domain use for embedding (DCT values) or the one in which it is easier to model the image (spatial / pixels). \nOne thing that is sure is that until now, most steganalysis methods have been using pixels values, IMHO because DCT coefficients are more difficult to interpret / model / analysis. \nThis is true for \"classical\" steganalysis that uses features extraction + supervised learning.\nThis is even more for Deep Learning based staganalysis: applying local convolution over DCT values seems very difficult becase DCT values next to each other are very different.... (hence the proposal in <a href=\"http://dde.binghamton.edu/vholub/pdf/Holub_PhD_Dissertation_2014.pdf\">chapitre 8 of the thesis you cited</a> to use of filters of the DCT without subsampling ....).</p>\n\n<p>I am also very surprised that even without separating image compressed with different quality factors (QF) pretty good results can be obtained ... or perhaps the top scorer are separating those image and train three different networks .... </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 835462,
          "author_name": "robga",
          "author_url": "",
          "post_date": "05/06/2020 09:28:28",
          "content": "<p>It is possible a network fed with pixels is learning an approximation of the DCT coefficients, and mixed QFs may provide regularisation. I am yet to try it, but stride 8 convolutions in the first layer of an architecture may provide some hinting to counter the fact that DCT values next to each other are very different.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "828180": "Hello There,\n\nI have seen many of you looking at the difference between images to \"see\" the difference between cover and stego image. \nI just wanted to state loudly that this is a consequence of data hiding. Indeed, the image are compressed in JPEG and you not be able to embed data in pixel ... after compression it would be unreadable.\nTo answer a question from @vovanf98 I made a small notebook (I apologize, I am terrible at programming) available here, *for which I deeply thank @yousfi for his help using external library*:\n[https://www.kaggle.com/remicogranne/inspect-impact-of-steganography-on-dct-coefs](https://www.kaggle.com/remicogranne/inspect-impact-of-steganography-on-dct-coefs)\nwhich aims to show several basic things:\n(1) data are hidden into DCT coefficients. Looking at pixels is a consequence of data hiding not the hidden data themselves.\n(2) Looking at pixels difference can give you the taste that difference (hidden data) are more or less evenly spread in all color. It is not ! Much more data are hidden in the Y channel of JPEG. Yet you can hardly see this in spatial domain.\n(3) Of course the embedding is adaptive in the sense that it embeds mostly around \"complex areas\", more bits in \"textured images\" and more bits in images compressed with higher JPEG quality factors.\n\n(4) Eventually, I would like share a few question that remain open in the community:\n- Should we do the detection in pixels or in DCT coefficients ? Pixels are more easy to model while the hidden data you want to detect really is hidden in DCT ....\n- We used three different JPEG quality factors (95/90/75) and three different embedding scheme. Is it better to merge everything or should you try to do the learning for each QF and each embedding (on testing you can always identify QF by looking at Qtable and use a fusion function for the embedding...)\n- Similarly, should you merge all images or try to separate them by content ? Noise level ? \nIn other words the best may be trying to be as universal as possible or as targeted as possible ...",
    "833318": "The answer to the question whether to look in the DCT coefficients or in the pixels for a pattern seems to be not so straightforward to answer. Many papers claim that it is common sense to use the same domain that is used for embedding. However, for example in this thesis (http://dde.binghamton.edu/vholub/pdf/Holub_PhD_Dissertation_2014.pdf) you can find a different conclusion (last chapter). It is claimed that also for DCT embedding it it is easier to detect the embedding in the spatial domain, since most algorithms use a cost function defined in the spatial domain.",
    "835301": "Hello there,\nI agree with your post, different source will provide different argument whether one should use the domain use for embedding (DCT values) or the one in which it is easier to model the image (spatial / pixels). \nOne thing that is sure is that until now, most steganalysis methods have been using pixels values, IMHO because DCT coefficients are more difficult to interpret / model / analysis. \nThis is true for \"classical\" steganalysis that uses features extraction + supervised learning.\nThis is even more for Deep Learning based staganalysis: applying local convolution over DCT values seems very difficult becase DCT values next to each other are very different.... (hence the proposal in [chapitre 8 of the thesis you cited](http://dde.binghamton.edu/vholub/pdf/Holub_PhD_Dissertation_2014.pdf) to use of filters of the DCT without subsampling ....).\n\nI am also very surprised that even without separating image compressed with different quality factors (QF) pretty good results can be obtained ... or perhaps the top scorer are separating those image and train three different networks ....",
    "835462": "It is possible a network fed with pixels is learning an approximation of the DCT coefficients, and mixed QFs may provide regularisation. I am yet to try it, but stride 8 convolutions in the first layer of an architecture may provide some hinting to counter the fact that DCT values next to each other are very different."
  },
  "source": "meta"
}