{
  "id": 167652,
  "title": "Model Error Analysis",
  "url": "/competitions/alaska2-image-steganalysis/discussion/167652",
  "author_name": "",
  "post_date": "2020-07-17T12:00:52.005293700Z",
  "votes": 12,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Normally after I set my validation, I go iterative starting from a basic model. I continuously check in what kind of examples model makes mistakes and try to improve on that. But in this competition, I can't visually see the difference between cover and stego images. When the model makes a mistake on validation set, I can only have some empathy, nothing else:) Do you experience the same? </p>",
  "messages": [
    {
      "id": "932972",
      "postDate": "07/17/2020 12:00:52",
      "content": "<p>Normally after I set my validation, I go iterative starting from a basic model. I continuously check in what kind of examples model makes mistakes and try to improve on that. But in this competition, I can't visually see the difference between cover and stego images. When the model makes a mistake on validation set, I can only have some empathy, nothing else:) Do you experience the same? </p>",
      "rawMarkdown": "Normally after I set my validation, I go iterative starting from a basic model. I continuously check in what kind of examples model makes mistakes and try to improve on that. But in this competition, I can't visually see the difference between cover and stego images. When the model makes a mistake on validation set, I can only have some empathy, nothing else:) Do you experience the same?",
      "votes": null
    },
    {
      "id": "933046",
      "postDate": "07/17/2020 12:58:53",
      "content": "<p>I experienced the same. This dataset is unique. </p>",
      "rawMarkdown": "I experienced the same. This dataset is unique.",
      "votes": null
    },
    {
      "id": "933111",
      "postDate": "07/17/2020 13:50:59",
      "content": "<p>Same here. \n<code>.. can't visually see the difference between cover and stego images</code> - damn true. 😯 \nI think the problem is kind of, the payload is hidden in DCT, and we're using pre-trained model that generalizes on pixel values. </p>",
      "rawMarkdown": "Same here. \n`.. can't visually see the difference between cover and stego images` - damn true. 😯 \nI think the problem is kind of, the payload is hidden in DCT, and we're using pre-trained model that generalizes on pixel values.",
      "votes": null
    },
    {
      "id": "933122",
      "postDate": "07/17/2020 14:01:48",
      "content": "<p>\"I think the problem is kind of, the payload is hidden in DCT, and we're using pre-trained model that generalizes on pixel values.\"</p>\n\n<p>I've been wondering if any of teams trained imagenet from scratch with dct as input, on efficientnet. And used that as pretraining for this task. So no pixelwise generalization involved. Alternatively, train imagenet from scratch on ycbcr, I was surprised to find little publically in this area of alternate imagenet colorspaces. The rules allow it. Not long to find out ...</p>",
      "rawMarkdown": "\"I think the problem is kind of, the payload is hidden in DCT, and we're using pre-trained model that generalizes on pixel values.\"\n\nI've been wondering if any of teams trained imagenet from scratch with dct as input, on efficientnet. And used that as pretraining for this task. So no pixelwise generalization involved. Alternatively, train imagenet from scratch on ycbcr, I was surprised to find little publically in this area of alternate imagenet colorspaces. The rules allow it. Not long to find out ...",
      "votes": null
    },
    {
      "id": "933147",
      "postDate": "07/17/2020 14:24:39",
      "content": "<p>I experienced the same. I look forward to hearing after the competition how others are doing it.</p>",
      "rawMarkdown": "I experienced the same. I look forward to hearing after the competition how others are doing it.",
      "votes": null
    },
    {
      "id": "933192",
      "postDate": "07/17/2020 14:57:27",
      "content": "<p>That would be great. But I believe that would be also a pure research work.</p>",
      "rawMarkdown": "That would be great. But I believe that would be also a pure research work.",
      "votes": null
    },
    {
      "id": "933196",
      "postDate": "07/17/2020 15:00:35",
      "content": "<p>You may not be able to visualize the differences very well in the image space, but one way to gain insight is to project the embeddings from the last layer of your network using UMAP or tSNE and look for where your models are having issues discriminating along class boundaries.  It should give some clues as to what's going on.  Here's an example from a high performing model.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/933196/16355/umap_projection_.jpg\" alt=\"\"></p>",
      "rawMarkdown": "You may not be able to visualize the differences very well in the image space, but one way to gain insight is to project the embeddings from the last layer of your network using UMAP or tSNE and look for where your models are having issues discriminating along class boundaries.  It should give some clues as to what's going on.  Here's an example from a high performing model.\n\n![](https://storage.googleapis.com/kaggle-forum-message-attachments/933196/16355/umap_projection_.jpg)",
      "votes": null
    },
    {
      "id": "933199",
      "postDate": "07/17/2020 15:01:04",
      "content": "<blockquote>\n  <p>I've been wondering if any of teams trained imagenet from scratch on efficientnet using dct or ycbcr as input. </p>\n</blockquote>\n\n<p>Now that's an idea.</p>",
      "rawMarkdown": "&gt; I've been wondering if any of teams trained imagenet from scratch on efficientnet using dct or ycbcr as input. \n\nNow that's an idea.",
      "votes": null
    },
    {
      "id": "933206",
      "postDate": "07/17/2020 15:06:02",
      "content": "<p>Great visualization and this give a new perspective to others.\nthanks for sharing it.\nAwesome! 👍 </p>",
      "rawMarkdown": "Great visualization and this give a new perspective to others.\nthanks for sharing it.\nAwesome! 👍",
      "votes": null
    },
    {
      "id": "935351",
      "postDate": "07/19/2020 10:03:39",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F120121c004d2c3fd630aab652635db27%2FSelection_037.png?generation=1595152988715693&amp;alt=media\" alt=\"\"></p>\n\n<p>you can make difference images like this to inspect what kinds of micro-pattern are detected or missed</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F120121c004d2c3fd630aab652635db27%2FSelection_037.png?generation=1595152988715693&amp;alt=media)\n\n\nyou can make difference images like this to inspect what kinds of micro-pattern are detected or missed",
      "votes": null
    },
    {
      "id": "935931",
      "postDate": "07/19/2020 19:33:53",
      "content": "<p>If you can visually see the difference between cover and stego images - in a few hours you can manually label all test images🙄</p>",
      "rawMarkdown": "If you can visually see the difference between cover and stego images - in a few hours you can manually label all test images🙄",
      "votes": null
    },
    {
      "id": "937323",
      "postDate": "07/20/2020 23:44:08",
      "content": "<p>Nice trick. I did the same thing in Melanoma Comp to compare the training images and test images distribution. Also when using external data, you can compare how well external image data overlaps train and test <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028</a></p>\n<p>Using UMAP or t-SNE is a powerful tool in image comps that often goes overlooked. Now with RAPIDS cuML, you can perform UMAP and t-SNE in seconds on large datasets!</p>",
      "rawMarkdown": "Nice trick. I did the same thing in Melanoma Comp to compare the training images and test images distribution. Also when using external data, you can compare how well external image data overlaps train and test https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028\n\nUsing UMAP or t-SNE is a powerful tool in image comps that often goes overlooked. Now with RAPIDS cuML, you can perform UMAP and t-SNE in seconds on large datasets!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 935931,
      "author_name": "golubev",
      "author_url": "",
      "post_date": "07/19/2020 19:33:53",
      "content": "<p>If you can visually see the difference between cover and stego images - in a few hours you can manually label all test images🙄</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 933046,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "07/17/2020 12:58:53",
      "content": "<p>I experienced the same. This dataset is unique. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 933111,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "07/17/2020 13:50:59",
      "content": "<p>Same here. \n<code>.. can't visually see the difference between cover and stego images</code> - damn true. 😯 \nI think the problem is kind of, the payload is hidden in DCT, and we're using pre-trained model that generalizes on pixel values. </p>",
      "votes": null,
      "replies": [
        {
          "id": 933122,
          "author_name": "robga",
          "author_url": "",
          "post_date": "07/17/2020 14:01:48",
          "content": "<p>\"I think the problem is kind of, the payload is hidden in DCT, and we're using pre-trained model that generalizes on pixel values.\"</p>\n\n<p>I've been wondering if any of teams trained imagenet from scratch with dct as input, on efficientnet. And used that as pretraining for this task. So no pixelwise generalization involved. Alternatively, train imagenet from scratch on ycbcr, I was surprised to find little publically in this area of alternate imagenet colorspaces. The rules allow it. Not long to find out ...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 933192,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "07/17/2020 14:57:27",
          "content": "<p>That would be great. But I believe that would be also a pure research work.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 933199,
          "author_name": "authman",
          "author_url": "",
          "post_date": "07/17/2020 15:01:04",
          "content": "<blockquote>\n  <p>I've been wondering if any of teams trained imagenet from scratch on efficientnet using dct or ycbcr as input. </p>\n</blockquote>\n\n<p>Now that's an idea.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 933147,
      "author_name": "zaburo",
      "author_url": "",
      "post_date": "07/17/2020 14:24:39",
      "content": "<p>I experienced the same. I look forward to hearing after the competition how others are doing it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 933196,
      "author_name": "tivfrvqhs5",
      "author_url": "",
      "post_date": "07/17/2020 15:00:35",
      "content": "<p>You may not be able to visualize the differences very well in the image space, but one way to gain insight is to project the embeddings from the last layer of your network using UMAP or tSNE and look for where your models are having issues discriminating along class boundaries.  It should give some clues as to what's going on.  Here's an example from a high performing model.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/933196/16355/umap_projection_.jpg\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 933206,
          "author_name": "piantic",
          "author_url": "",
          "post_date": "07/17/2020 15:06:02",
          "content": "<p>Great visualization and this give a new perspective to others.\nthanks for sharing it.\nAwesome! 👍 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 937323,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/20/2020 23:44:08",
          "content": "<p>Nice trick. I did the same thing in Melanoma Comp to compare the training images and test images distribution. Also when using external data, you can compare how well external image data overlaps train and test <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028</a></p>\n<p>Using UMAP or t-SNE is a powerful tool in image comps that often goes overlooked. Now with RAPIDS cuML, you can perform UMAP and t-SNE in seconds on large datasets!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 935351,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/19/2020 10:03:39",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F120121c004d2c3fd630aab652635db27%2FSelection_037.png?generation=1595152988715693&amp;alt=media\" alt=\"\"></p>\n\n<p>you can make difference images like this to inspect what kinds of micro-pattern are detected or missed</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "932972": "Normally after I set my validation, I go iterative starting from a basic model. I continuously check in what kind of examples model makes mistakes and try to improve on that. But in this competition, I can't visually see the difference between cover and stego images. When the model makes a mistake on validation set, I can only have some empathy, nothing else:) Do you experience the same?",
    "933046": "I experienced the same. This dataset is unique.",
    "933111": "Same here. \n`.. can't visually see the difference between cover and stego images` - damn true. 😯 \nI think the problem is kind of, the payload is hidden in DCT, and we're using pre-trained model that generalizes on pixel values.",
    "933122": "\"I think the problem is kind of, the payload is hidden in DCT, and we're using pre-trained model that generalizes on pixel values.\"\n\nI've been wondering if any of teams trained imagenet from scratch with dct as input, on efficientnet. And used that as pretraining for this task. So no pixelwise generalization involved. Alternatively, train imagenet from scratch on ycbcr, I was surprised to find little publically in this area of alternate imagenet colorspaces. The rules allow it. Not long to find out ...",
    "933147": "I experienced the same. I look forward to hearing after the competition how others are doing it.",
    "933192": "That would be great. But I believe that would be also a pure research work.",
    "933196": "You may not be able to visualize the differences very well in the image space, but one way to gain insight is to project the embeddings from the last layer of your network using UMAP or tSNE and look for where your models are having issues discriminating along class boundaries.  It should give some clues as to what's going on.  Here's an example from a high performing model.\n\n![](https://storage.googleapis.com/kaggle-forum-message-attachments/933196/16355/umap_projection_.jpg)",
    "933199": "&gt; I've been wondering if any of teams trained imagenet from scratch on efficientnet using dct or ycbcr as input. \n\nNow that's an idea.",
    "933206": "Great visualization and this give a new perspective to others.\nthanks for sharing it.\nAwesome! 👍",
    "935351": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F120121c004d2c3fd630aab652635db27%2FSelection_037.png?generation=1595152988715693&amp;alt=media)\n\n\nyou can make difference images like this to inspect what kinds of micro-pattern are detected or missed",
    "935931": "If you can visually see the difference between cover and stego images - in a few hours you can manually label all test images🙄",
    "937323": "Nice trick. I did the same thing in Melanoma Comp to compare the training images and test images distribution. Also when using external data, you can compare how well external image data overlaps train and test https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028\n\nUsing UMAP or t-SNE is a powerful tool in image comps that often goes overlooked. Now with RAPIDS cuML, you can perform UMAP and t-SNE in seconds on large datasets!"
  },
  "source": "meta"
}