{
  "id": 165395,
  "title": "A Return to DCT",
  "url": "/competitions/alaska2-image-steganalysis/discussion/165395",
  "author_name": "",
  "post_date": "2020-07-09T14:53:03.327351900Z",
  "votes": 13,
  "comment_count": 6,
  "views": 0,
  "content": "<p>As with most of you, I have a lot of down time while waiting for models to converge up. I spent some of this time yesterday coming up with a few new ideas and even build out pipelines to test a couple of them. One of the more interesting ones, in my opinion, builds off of work done by Uber AI Lab's JPEG2DCT library, which supposedly speeds up network training time by cutting the full JPEG -&gt; DCT -&gt; YCbCr -&gt; RGB transformation pipeline in half and feeding DCT components directly into an FCN. They were able to achieve SoTA results doing so in 2018. I've prepared <a href=\"https://www.kaggle.com/authman/deep-explorations-with-ubers-jpeg2dct\">this kernel</a> to discuss my findings..</p>",
  "messages": [
    {
      "id": "921786",
      "postDate": "07/09/2020 14:53:03",
      "content": "<p>As with most of you, I have a lot of down time while waiting for models to converge up. I spent some of this time yesterday coming up with a few new ideas and even build out pipelines to test a couple of them. One of the more interesting ones, in my opinion, builds off of work done by Uber AI Lab's JPEG2DCT library, which supposedly speeds up network training time by cutting the full JPEG -&gt; DCT -&gt; YCbCr -&gt; RGB transformation pipeline in half and feeding DCT components directly into an FCN. They were able to achieve SoTA results doing so in 2018. I've prepared <a href=\"https://www.kaggle.com/authman/deep-explorations-with-ubers-jpeg2dct\">this kernel</a> to discuss my findings..</p>",
      "rawMarkdown": "As with most of you, I have a lot of down time while waiting for models to converge up. I spent some of this time yesterday coming up with a few new ideas and even build out pipelines to test a couple of them. One of the more interesting ones, in my opinion, builds off of work done by Uber AI Lab's JPEG2DCT library, which supposedly speeds up network training time by cutting the full JPEG -&gt; DCT -&gt; YCbCr -&gt; RGB transformation pipeline in half and feeding DCT components directly into an FCN. They were able to achieve SoTA results doing so in 2018. I've prepared [this kernel](https://www.kaggle.com/authman/deep-explorations-with-ubers-jpeg2dct) to discuss my findings..",
      "votes": null
    },
    {
      "id": "921816",
      "postDate": "07/09/2020 15:22:28",
      "content": "<p>Thanks <a href=\"/authman\">@authman</a> for the  time and efforts you put in preparing and sharing the kernel with the community. Will try this approach also. </p>",
      "rawMarkdown": "Thanks @authman for the  time and efforts you put in preparing and sharing the kernel with the community. Will try this approach also.",
      "votes": null
    },
    {
      "id": "922003",
      "postDate": "07/09/2020 18:11:46",
      "content": "<p>Much appreciated effort! Thanks!</p>",
      "rawMarkdown": "Much appreciated effort! Thanks!",
      "votes": null
    },
    {
      "id": "923784",
      "postDate": "07/11/2020 06:10:14",
      "content": "<p>Great! I spent over a month trying to make the DCT domain work, but I never got close to a competitive score. However, I lack the deep learning experience and hardware. I still cannot fathom why detection in the spatial domain would be better than in the DCT domain, where the embedding changes are made. The only reason I can think of are very well trained existing spatial domain models.\nI am curious about the results!</p>\n\n<p>Vojtech Holub, creator of the JUNIWARD</p>",
      "rawMarkdown": "Great! I spent over a month trying to make the DCT domain work, but I never got close to a competitive score. However, I lack the deep learning experience and hardware. I still cannot fathom why detection in the spatial domain would be better than in the DCT domain, where the embedding changes are made. The only reason I can think of are very well trained existing spatial domain models.\nI am curious about the results!\n\nVojtech Holub, creator of the JUNIWARD",
      "votes": null
    },
    {
      "id": "924655",
      "postDate": "07/11/2020 15:08:45",
      "content": "<p>Good to see you here, <a href=\"/vojtechholub\">@vojtechholub</a> ! Yes making a good JUNI detector in the DCT domain is not easy... Maybe your results from SPIE 2014 are still true with deep learning based detectors ;) </p>\n\n<blockquote>\n  <p>[...] we demonstrate that more accurate detection is obtained when constructing the steganalysis features in the spatial domain where the distortion function is minimized, challenging thus both established doctrines. \n  Challenging the Doctrines of JPEG Steganography, Vojtěch Holub and Jessica Fridrich. </p>\n</blockquote>",
      "rawMarkdown": "Good to see you here, @vojtechholub ! Yes making a good JUNI detector in the DCT domain is not easy... Maybe your results from SPIE 2014 are still true with deep learning based detectors ;) \n&gt; [...] we demonstrate that more accurate detection is obtained when constructing the steganalysis features in the spatial domain where the distortion function is minimized, challenging thus both established doctrines. \nChallenging the Doctrines of JPEG Steganography, Vojtěch Holub and Jessica Fridrich.",
      "votes": null
    },
    {
      "id": "924907",
      "postDate": "07/11/2020 17:35:00",
      "content": "<p>Yassine, I hoped that nobody would use that against me :)\nMy thinking at that time was that we have no idea how to well model DCT/JPEG domain images, but we were pretty good at modeling spatial domain.</p>\n\n<p>Now we have a tool, DNN, that should allow us to model DCT much better.</p>\n\n<p>Now, correct me if I'm wrong, but the transformation between quantized DCT and non-rounded spatial domain is purely linear. That is something that DNN should learn easily and I still don't understand why learning one domain is more difficult to learn then the other linearly dependent domain.</p>",
      "rawMarkdown": "Yassine, I hoped that nobody would use that against me :)\nMy thinking at that time was that we have no idea how to well model DCT/JPEG domain images, but we were pretty good at modeling spatial domain.\n\nNow we have a tool, DNN, that should allow us to model DCT much better.\n\nNow, correct me if I'm wrong, but the transformation between quantized DCT and non-rounded spatial domain is purely linear. That is something that DNN should learn easily and I still don't understand why learning one domain is more difficult to learn then the other linearly dependent domain.",
      "votes": null
    },
    {
      "id": "925312",
      "postDate": "07/12/2020 02:23:15",
      "content": "<p>Sorry to double post <a href=\"https://www.kaggle.com/authman/deep-explorations-with-ubers-jpeg2dct/comments#925309\">this</a>, but then I thought the discussion forums are where the discussions happen 😉 </p>\n\n<p>This can be easily verified. You're correct in saying that models benefit from the pre-training. These pre-trained weights are used because Deep Learning models require humongous amounts of data to converge, which in practice, is a rarity. And to train with that amount of data requires some serious hardware (with I think you guys have, and I know I don't have 😆). When using pre-trained weights, the models converge with less amount of data and in fewer epochs because they already have a general idea of how the world looks. Courtesy, Transfer Learning.</p>\n\n<p>IMO here, 300k images are quite sufficient to drive a model to convergence. If not, you can just generate more data.\nWhat you can do is, take a well-performing model, such as resnet50 or one of the efficientnets and train them in the DCT domain from scratch, i.e., instead of starting with the pre-trained weights start with random weights. Let it train for 300 epochs. That'll be your baseline to improve upon.</p>\n\n<p>Eager to know the results, if you try this out,\nCheers!</p>",
      "rawMarkdown": "Sorry to double post [this](https://www.kaggle.com/authman/deep-explorations-with-ubers-jpeg2dct/comments#925309), but then I thought the discussion forums are where the discussions happen 😉 \n\nThis can be easily verified. You're correct in saying that models benefit from the pre-training. These pre-trained weights are used because Deep Learning models require humongous amounts of data to converge, which in practice, is a rarity. And to train with that amount of data requires some serious hardware (with I think you guys have, and I know I don't have 😆). When using pre-trained weights, the models converge with less amount of data and in fewer epochs because they already have a general idea of how the world looks. Courtesy, Transfer Learning.\n\nIMO here, 300k images are quite sufficient to drive a model to convergence. If not, you can just generate more data.\nWhat you can do is, take a well-performing model, such as resnet50 or one of the efficientnets and train them in the DCT domain from scratch, i.e., instead of starting with the pre-trained weights start with random weights. Let it train for 300 epochs. That'll be your baseline to improve upon.\n\nEager to know the results, if you try this out,\nCheers!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 921816,
      "author_name": "vishnurapps",
      "author_url": "",
      "post_date": "07/09/2020 15:22:28",
      "content": "<p>Thanks <a href=\"/authman\">@authman</a> for the  time and efforts you put in preparing and sharing the kernel with the community. Will try this approach also. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 922003,
      "author_name": "mightyrains",
      "author_url": "",
      "post_date": "07/09/2020 18:11:46",
      "content": "<p>Much appreciated effort! Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 923784,
      "author_name": "vojtechholub",
      "author_url": "",
      "post_date": "07/11/2020 06:10:14",
      "content": "<p>Great! I spent over a month trying to make the DCT domain work, but I never got close to a competitive score. However, I lack the deep learning experience and hardware. I still cannot fathom why detection in the spatial domain would be better than in the DCT domain, where the embedding changes are made. The only reason I can think of are very well trained existing spatial domain models.\nI am curious about the results!</p>\n\n<p>Vojtech Holub, creator of the JUNIWARD</p>",
      "votes": null,
      "replies": [
        {
          "id": 924655,
          "author_name": "yousfi",
          "author_url": "",
          "post_date": "07/11/2020 15:08:45",
          "content": "<p>Good to see you here, <a href=\"/vojtechholub\">@vojtechholub</a> ! Yes making a good JUNI detector in the DCT domain is not easy... Maybe your results from SPIE 2014 are still true with deep learning based detectors ;) </p>\n\n<blockquote>\n  <p>[...] we demonstrate that more accurate detection is obtained when constructing the steganalysis features in the spatial domain where the distortion function is minimized, challenging thus both established doctrines. \n  Challenging the Doctrines of JPEG Steganography, Vojtěch Holub and Jessica Fridrich. </p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 924907,
          "author_name": "vojtechholub",
          "author_url": "",
          "post_date": "07/11/2020 17:35:00",
          "content": "<p>Yassine, I hoped that nobody would use that against me :)\nMy thinking at that time was that we have no idea how to well model DCT/JPEG domain images, but we were pretty good at modeling spatial domain.</p>\n\n<p>Now we have a tool, DNN, that should allow us to model DCT much better.</p>\n\n<p>Now, correct me if I'm wrong, but the transformation between quantized DCT and non-rounded spatial domain is purely linear. That is something that DNN should learn easily and I still don't understand why learning one domain is more difficult to learn then the other linearly dependent domain.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 925312,
          "author_name": "mightyrains",
          "author_url": "",
          "post_date": "07/12/2020 02:23:15",
          "content": "<p>Sorry to double post <a href=\"https://www.kaggle.com/authman/deep-explorations-with-ubers-jpeg2dct/comments#925309\">this</a>, but then I thought the discussion forums are where the discussions happen 😉 </p>\n\n<p>This can be easily verified. You're correct in saying that models benefit from the pre-training. These pre-trained weights are used because Deep Learning models require humongous amounts of data to converge, which in practice, is a rarity. And to train with that amount of data requires some serious hardware (with I think you guys have, and I know I don't have 😆). When using pre-trained weights, the models converge with less amount of data and in fewer epochs because they already have a general idea of how the world looks. Courtesy, Transfer Learning.</p>\n\n<p>IMO here, 300k images are quite sufficient to drive a model to convergence. If not, you can just generate more data.\nWhat you can do is, take a well-performing model, such as resnet50 or one of the efficientnets and train them in the DCT domain from scratch, i.e., instead of starting with the pre-trained weights start with random weights. Let it train for 300 epochs. That'll be your baseline to improve upon.</p>\n\n<p>Eager to know the results, if you try this out,\nCheers!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "921786": "As with most of you, I have a lot of down time while waiting for models to converge up. I spent some of this time yesterday coming up with a few new ideas and even build out pipelines to test a couple of them. One of the more interesting ones, in my opinion, builds off of work done by Uber AI Lab's JPEG2DCT library, which supposedly speeds up network training time by cutting the full JPEG -&gt; DCT -&gt; YCbCr -&gt; RGB transformation pipeline in half and feeding DCT components directly into an FCN. They were able to achieve SoTA results doing so in 2018. I've prepared [this kernel](https://www.kaggle.com/authman/deep-explorations-with-ubers-jpeg2dct) to discuss my findings..",
    "921816": "Thanks @authman for the  time and efforts you put in preparing and sharing the kernel with the community. Will try this approach also.",
    "922003": "Much appreciated effort! Thanks!",
    "923784": "Great! I spent over a month trying to make the DCT domain work, but I never got close to a competitive score. However, I lack the deep learning experience and hardware. I still cannot fathom why detection in the spatial domain would be better than in the DCT domain, where the embedding changes are made. The only reason I can think of are very well trained existing spatial domain models.\nI am curious about the results!\n\nVojtech Holub, creator of the JUNIWARD",
    "924655": "Good to see you here, @vojtechholub ! Yes making a good JUNI detector in the DCT domain is not easy... Maybe your results from SPIE 2014 are still true with deep learning based detectors ;) \n&gt; [...] we demonstrate that more accurate detection is obtained when constructing the steganalysis features in the spatial domain where the distortion function is minimized, challenging thus both established doctrines. \nChallenging the Doctrines of JPEG Steganography, Vojtěch Holub and Jessica Fridrich.",
    "924907": "Yassine, I hoped that nobody would use that against me :)\nMy thinking at that time was that we have no idea how to well model DCT/JPEG domain images, but we were pretty good at modeling spatial domain.\n\nNow we have a tool, DNN, that should allow us to model DCT much better.\n\nNow, correct me if I'm wrong, but the transformation between quantized DCT and non-rounded spatial domain is purely linear. That is something that DNN should learn easily and I still don't understand why learning one domain is more difficult to learn then the other linearly dependent domain.",
    "925312": "Sorry to double post [this](https://www.kaggle.com/authman/deep-explorations-with-ubers-jpeg2dct/comments#925309), but then I thought the discussion forums are where the discussions happen 😉 \n\nThis can be easily verified. You're correct in saying that models benefit from the pre-training. These pre-trained weights are used because Deep Learning models require humongous amounts of data to converge, which in practice, is a rarity. And to train with that amount of data requires some serious hardware (with I think you guys have, and I know I don't have 😆). When using pre-trained weights, the models converge with less amount of data and in fewer epochs because they already have a general idea of how the world looks. Courtesy, Transfer Learning.\n\nIMO here, 300k images are quite sufficient to drive a model to convergence. If not, you can just generate more data.\nWhat you can do is, take a well-performing model, such as resnet50 or one of the efficientnets and train them in the DCT domain from scratch, i.e., instead of starting with the pre-trained weights start with random weights. Let it train for 300 epochs. That'll be your baseline to improve upon.\n\nEager to know the results, if you try this out,\nCheers!"
  },
  "source": "meta"
}