{
  "id": 166331,
  "title": "Single vs Ensemble Scores.",
  "url": "/competitions/alaska2-image-steganalysis/discussion/166331",
  "author_name": "Innat",
  "post_date": "2020-07-12T14:19:43.282000",
  "votes": 2,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I've two submissions from two different models, one of them scored <code>0.910</code> another <code>0.911</code>.  I've tried to blend these two but got worse results, ~80. I'm wondering if anyone is facing this? As the target images are sensitive, I'm wondering do we need to consider some specialty here. </p>\n\n<p>However, the ensemble with the same architectural models sometimes gave a reasonable score but less. Also in my case, TTA improves a single model but worse in an ensemble.  </p>\n\n<h3>Update</h3>\n\n<p><strong>Solved</strong></p>",
  "messages": [
    {
      "id": 934202,
      "postDate": "2020-07-18T09:59:53.220Z",
      "content": "<p>\"I've two submissions from two different models, one of them scored 0.910 another 0.911. I've tried to blend these two but got worse results, ~80.\"</p>\n\n<p>you can break the problems into subproblems.</p>\n\n<p>the test image has equal number of qf75,90,95 images.\nit is mentioned that the type of steganography happens with equal probability.</p>\n\n<p>we already know that validation accuracy on different subgroups (qf, steganography) are different.\nyou can do experiments to see if TTA will improves different subgroups or not.</p>\n\n<p>Then you can  decide if you want to TTA for the test samples subgroups or not since you know the qf (and approximate steganography type) of the test samples   </p>\n\n<hr>\n\n<p>there is one more information available to both test/train, but i haven't investigated yet: num of non-zero  ac coefficients.</p>",
      "rawMarkdown": "\"I've two submissions from two different models, one of them scored 0.910 another 0.911. I've tried to blend these two but got worse results, ~80.\"\n\n\nyou can break the problems into subproblems.\n\nthe test image has equal number of qf75,90,95 images.\nit is mentioned that the type of steganography happens with equal probability.\n\nwe already know that validation accuracy on different subgroups (qf, steganography) are different.\nyou can do experiments to see if TTA will improves different subgroups or not.\n\nThen you can  decide if you want to TTA for the test samples subgroups or not since you know the qf (and approximate steganography type) of the test samples   \n\n\n\n---\n\nthere is one more information available to both test/train, but i haven't investigated yet: num of non-zero  ac coefficients.\n",
      "votes": 1
    },
    {
      "id": 926210,
      "postDate": "2020-07-12T15:20:12.907Z",
      "content": "<p>It is for sure a bug. While blending always make sure to join 2 submissions on ids.</p>",
      "rawMarkdown": "It is for sure a bug. While blending always make sure to join 2 submissions on ids.",
      "votes": 1,
      "replies": [
        {
          "id": 926281,
          "postDate": "2020-07-12T16:08:22.817Z",
          "content": "<p>Yes, thank you. I've checked the id though :(</p>",
          "rawMarkdown": "Yes, thank you. I've checked the id though :("
        }
      ]
    },
    {
      "id": 926361,
      "postDate": "2020-07-12T16:59:31.477Z",
      "content": "<p>also, try rank averaging </p>",
      "rawMarkdown": "also, try rank averaging ",
      "votes": 2,
      "replies": [
        {
          "id": 926647,
          "postDate": "2020-07-12T20:37:00.017Z",
          "content": "<p>thanks for the tips :)</p>",
          "rawMarkdown": "thanks for the tips :)"
        },
        {
          "id": 927910,
          "postDate": "2020-07-13T16:24:56.547Z",
          "content": "<p>Could you explain a little more? :)\n<a href=\"/yifanxie\">@yifanxie</a> </p>",
          "rawMarkdown": "Could you explain a little more? :)\n@yifanxie "
        },
        {
          "id": 928047,
          "postDate": "2020-07-13T18:00:43.273Z",
          "content": "<blockquote>\n  <p><strong>Heroseo wrote:</strong></p>\n  \n  <p>Could you explain a little more? :)\n  <a href=\"/yifanxie\">@yifanxie</a> </p>\n</blockquote>\n\n<p>For AUC related metric, only the rank of the probability matters - so you can convert probability predictions to ranking using functions like Panda's <a href=\"https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.rank.html\">DataFrame.rank</a>, and then merge the ranks of predictions from different models.</p>\n\n<p>Some models tend to generate probability values of very different magnitude even within the same [0,1] interval, and as a result, there are situations in which the average prediction get dominated by probability output of larger magnitute.  In such situations,  doing rank averaging operation remove the side effect of merging by raw probability values </p>",
          "rawMarkdown": "&gt; **Heroseo wrote:**\n&gt; \n&gt; Could you explain a little more? :)\n&gt; @yifanxie \n\nFor AUC related metric, only the rank of the probability matters - so you can convert probability predictions to ranking using functions like Panda's [DataFrame.rank](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.rank.html), and then merge the ranks of predictions from different models.\n\nSome models tend to generate probability values of very different magnitude even within the same [0,1] interval, and as a result, there are situations in which the average prediction get dominated by probability output of larger magnitute.  In such situations,  doing rank averaging operation remove the side effect of merging by raw probability values \n",
          "votes": 4
        },
        {
          "id": 928105,
          "postDate": "2020-07-13T18:53:18.897Z",
          "content": "<p>Thanks a lot! :)\n<a href=\"/yifanxie\">@yifanxie</a> </p>",
          "rawMarkdown": "Thanks a lot! :)\n@yifanxie "
        }
      ]
    },
    {
      "id": 926135,
      "postDate": "2020-07-12T14:19:43.283Z",
      "content": "<p>I've two submissions from two different models, one of them scored <code>0.910</code> another <code>0.911</code>.  I've tried to blend these two but got worse results, ~80. I'm wondering if anyone is facing this? As the target images are sensitive, I'm wondering do we need to consider some specialty here. </p>\n\n<p>However, the ensemble with the same architectural models sometimes gave a reasonable score but less. Also in my case, TTA improves a single model but worse in an ensemble.  </p>\n\n<h3>Update</h3>\n\n<p><strong>Solved</strong></p>",
      "rawMarkdown": "I've two submissions from two different models, one of them scored `0.910` another `0.911`.  I've tried to blend these two but got worse results, ~80. I'm wondering if anyone is facing this? As the target images are sensitive, I'm wondering do we need to consider some specialty here. \n\nHowever, the ensemble with the same architectural models sometimes gave a reasonable score but less. Also in my case, TTA improves a single model but worse in an ensemble.  \n\n### Update\n**Solved**",
      "votes": 2
    },
    {
      "id": 928045,
      "postDate": "2020-07-13T17:56:48.140Z",
      "content": "<p>Just you must reset the index and make the ids sorted then you can do your blending easily !</p>",
      "rawMarkdown": "Just you must reset the index and make the ids sorted then you can do your blending easily !",
      "replies": [
        {
          "id": 932127,
          "postDate": "2020-07-16T18:17:39.773Z",
          "content": "<p>Yes, I'm aware of it. Thank you.</p>",
          "rawMarkdown": "Yes, I'm aware of it. Thank you."
        }
      ]
    },
    {
      "id": 935264,
      "postDate": "2020-07-19T07:58:41.587Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 934202,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-07-18T09:59:53.220000",
      "content": "<p>\"I've two submissions from two different models, one of them scored 0.910 another 0.911. I've tried to blend these two but got worse results, ~80.\"</p>\n\n<p>you can break the problems into subproblems.</p>\n\n<p>the test image has equal number of qf75,90,95 images.\nit is mentioned that the type of steganography happens with equal probability.</p>\n\n<p>we already know that validation accuracy on different subgroups (qf, steganography) are different.\nyou can do experiments to see if TTA will improves different subgroups or not.</p>\n\n<p>Then you can  decide if you want to TTA for the test samples subgroups or not since you know the qf (and approximate steganography type) of the test samples   </p>\n\n<hr>\n\n<p>there is one more information available to both test/train, but i haven't investigated yet: num of non-zero  ac coefficients.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 926210,
      "author_name": "Ahmet Erdem",
      "author_url": "",
      "post_date": "2020-07-12T15:20:12.907000",
      "content": "<p>It is for sure a bug. While blending always make sure to join 2 submissions on ids.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 926281,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-07-12T16:08:22.817000",
          "content": "<p>Yes, thank you. I've checked the id though :(</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 926361,
      "author_name": "Yifan Xie",
      "author_url": "",
      "post_date": "2020-07-12T16:59:31.477000",
      "content": "<p>also, try rank averaging </p>",
      "votes": 2,
      "replies": [
        {
          "id": 926647,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-07-12T20:37:00.017000",
          "content": "<p>thanks for the tips :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 927910,
          "author_name": "Heroseo",
          "author_url": "",
          "post_date": "2020-07-13T16:24:56.547000",
          "content": "<p>Could you explain a little more? :)\n<a href=\"/yifanxie\">@yifanxie</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 928047,
          "author_name": "Yifan Xie",
          "author_url": "",
          "post_date": "2020-07-13T18:00:43.273000",
          "content": "<blockquote>\n  <p><strong>Heroseo wrote:</strong></p>\n  \n  <p>Could you explain a little more? :)\n  <a href=\"/yifanxie\">@yifanxie</a> </p>\n</blockquote>\n\n<p>For AUC related metric, only the rank of the probability matters - so you can convert probability predictions to ranking using functions like Panda's <a href=\"https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.rank.html\">DataFrame.rank</a>, and then merge the ranks of predictions from different models.</p>\n\n<p>Some models tend to generate probability values of very different magnitude even within the same [0,1] interval, and as a result, there are situations in which the average prediction get dominated by probability output of larger magnitute.  In such situations,  doing rank averaging operation remove the side effect of merging by raw probability values </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 928105,
          "author_name": "Heroseo",
          "author_url": "",
          "post_date": "2020-07-13T18:53:18.897000",
          "content": "<p>Thanks a lot! :)\n<a href=\"/yifanxie\">@yifanxie</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 928045,
      "author_name": "Ashraf Mahdhi",
      "author_url": "",
      "post_date": "2020-07-13T17:56:48.140000",
      "content": "<p>Just you must reset the index and make the ids sorted then you can do your blending easily !</p>",
      "votes": 0,
      "replies": [
        {
          "id": 932127,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-07-16T18:17:39.773000",
          "content": "<p>Yes, I'm aware of it. Thank you.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 935264,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-19T07:58:41.587000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "934202": "\"I've two submissions from two different models, one of them scored 0.910 another 0.911. I've tried to blend these two but got worse results, ~80.\"\n\n\nyou can break the problems into subproblems.\n\nthe test image has equal number of qf75,90,95 images.\nit is mentioned that the type of steganography happens with equal probability.\n\nwe already know that validation accuracy on different subgroups (qf, steganography) are different.\nyou can do experiments to see if TTA will improves different subgroups or not.\n\nThen you can  decide if you want to TTA for the test samples subgroups or not since you know the qf (and approximate steganography type) of the test samples   \n\n\n\n---\n\nthere is one more information available to both test/train, but i haven't investigated yet: num of non-zero  ac coefficients.\n",
    "926210": "It is for sure a bug. While blending always make sure to join 2 submissions on ids.",
    "926361": "also, try rank averaging ",
    "926135": "I've two submissions from two different models, one of them scored `0.910` another `0.911`.  I've tried to blend these two but got worse results, ~80. I'm wondering if anyone is facing this? As the target images are sensitive, I'm wondering do we need to consider some specialty here. \n\nHowever, the ensemble with the same architectural models sometimes gave a reasonable score but less. Also in my case, TTA improves a single model but worse in an ensemble.  \n\n### Update\n**Solved**",
    "928045": "Just you must reset the index and make the ids sorted then you can do your blending easily !",
    "935264": ""
  }
}