{
  "id": 58377,
  "title": "Tuning final scores with several results, and about RMSE score",
  "url": "/competitions/avito-demand-prediction/discussion/58377",
  "author_name": "Ethan Sukhyun Hong",
  "post_date": "2018-06-07T01:30:00.914000",
  "votes": 9,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I found an interesting <a href=\"https://www.kaggle.com/moussaid/moyen-just-simple-moyen-of-my-predictions\">kernal :</a> \nThis one just try to find converging point between the two results - by repeat giving average of two results. For example, like this:</p>\n\n<p>import d1, d2 (two results)</p>\n\n<p>d1 = (d1+d2)/2</p>\n\n<p>d2 = (d1+d2)/2</p>\n\n<p>d1 = (d1+d2)/2</p>\n\n<p>d2 = (d1+d2)/2</p>\n\n<p>....</p>\n\n<p>I found that this method actually improves the RMSE score. I also tried this with more than one results, and interestingly it (almost everytime) gave improved results. </p>\n\n<p>Some may think this is just a trick, but I rather found it's statistically valid - since the scoring is based on RMSE, this method reduces variance within the results and reduce errors from outliers. So the result improves as much as the merging process reduces errors caused by variance itself. </p>\n\n<p>However, this raises me a question that 'is the result produced by this merging process could be seen as a <strong>Good prediction result</strong>?' And also this raised me a question of  'Is RMSE a good way to value preciseness of prediction result?'</p>\n\n<p>Low RMSE means in average, entities have low errors. It shows general tendencies of entire dataset. However, it doesn't count how much of the dataset got proper result, and how much haven't. </p>\n\n<p>Suppose that two predictions got same RMSE scores - One got entire rows incorrect, but to only little extent. The other got half of rows correct, but the other half incorrect to huge extent. Can we say these two prediction have identical validity in terms of correctness?</p>\n\n<p>(Anyway I believe this merging method could be a good way to tune your scores before submission for better scores!)</p>",
  "messages": [
    {
      "id": 340501,
      "postDate": "2018-06-09T11:26:13.963Z",
      "content": "<p>Hi,</p>\n\n<p>From my perspective, you are just performing a weighted average on your predictions and you are lucky to find good weights.</p>\n\n<p>Here is why: let's say that you load predictions from two files x and y:</p>\n\n<p>d1 = x</p>\n\n<p>d2 = y</p>\n\n<p>Then I apply your algorithm:</p>\n\n<p>d1 = (d1+d2)/2 = (x+y)/2 = 1/2*x + 1/2*y</p>\n\n<p>d2 = (d1+d2)/2 = (1/2*x + 1/2*y + y)/2 = 1/4*x + 3/4*y</p>\n\n<p>d1 = (d1+d2)/2 = (1/2*x + 1/2*y + 1/4*x + 3/4*y)/2 = 3/8*x + 5/8*y</p>\n\n<p>d2 = (d1+d2)/2 = (3/8*x + 5/8*y + 1/4*x + 3/4*y)/2 = 5/16*x + 11/16*y</p>\n\n<p>....</p>\n\n<p>Which is a simple weighted average of x and y with some \"magic\" weights.\nMoreover, simply exchanging x and y will give you complete different results.</p>",
      "rawMarkdown": "Hi,\n\nFrom my perspective, you are just performing a weighted average on your predictions and you are lucky to find good weights.\n\nHere is why: let's say that you load predictions from two files x and y:\n\nd1 = x\n\nd2 = y\n\nThen I apply your algorithm:\n\nd1 = (d1+d2)/2 = (x+y)/2 = 1/2*x + 1/2*y\n\nd2 = (d1+d2)/2 = (1/2*x + 1/2*y + y)/2 = 1/4*x + 3/4*y\n\nd1 = (d1+d2)/2 = (1/2*x + 1/2*y + 1/4*x + 3/4*y)/2 = 3/8*x + 5/8*y\n\nd2 = (d1+d2)/2 = (3/8*x + 5/8*y + 1/4*x + 3/4*y)/2 = 5/16*x + 11/16*y\n\n....\n\nWhich is a simple weighted average of x and y with some \"magic\" weights.\nMoreover, simply exchanging x and y will give you complete different results.",
      "votes": 12,
      "replies": [
        {
          "id": 340739,
          "postDate": "2018-06-10T07:21:22.563Z",
          "content": "<p>I think that's a good point thank you!</p>",
          "rawMarkdown": "I think that's a good point thank you!"
        }
      ]
    },
    {
      "id": 339448,
      "postDate": "2018-06-07T01:30:00.913Z",
      "content": "<p>I found an interesting <a href=\"https://www.kaggle.com/moussaid/moyen-just-simple-moyen-of-my-predictions\">kernal :</a> \nThis one just try to find converging point between the two results - by repeat giving average of two results. For example, like this:</p>\n\n<p>import d1, d2 (two results)</p>\n\n<p>d1 = (d1+d2)/2</p>\n\n<p>d2 = (d1+d2)/2</p>\n\n<p>d1 = (d1+d2)/2</p>\n\n<p>d2 = (d1+d2)/2</p>\n\n<p>....</p>\n\n<p>I found that this method actually improves the RMSE score. I also tried this with more than one results, and interestingly it (almost everytime) gave improved results. </p>\n\n<p>Some may think this is just a trick, but I rather found it's statistically valid - since the scoring is based on RMSE, this method reduces variance within the results and reduce errors from outliers. So the result improves as much as the merging process reduces errors caused by variance itself. </p>\n\n<p>However, this raises me a question that 'is the result produced by this merging process could be seen as a <strong>Good prediction result</strong>?' And also this raised me a question of  'Is RMSE a good way to value preciseness of prediction result?'</p>\n\n<p>Low RMSE means in average, entities have low errors. It shows general tendencies of entire dataset. However, it doesn't count how much of the dataset got proper result, and how much haven't. </p>\n\n<p>Suppose that two predictions got same RMSE scores - One got entire rows incorrect, but to only little extent. The other got half of rows correct, but the other half incorrect to huge extent. Can we say these two prediction have identical validity in terms of correctness?</p>\n\n<p>(Anyway I believe this merging method could be a good way to tune your scores before submission for better scores!)</p>",
      "rawMarkdown": "I found an interesting [kernal :][1] \nThis one just try to find converging point between the two results - by repeat giving average of two results. For example, like this:\n\nimport d1, d2 (two results)\n\nd1 = (d1+d2)/2\n\nd2 = (d1+d2)/2\n\nd1 = (d1+d2)/2\n\nd2 = (d1+d2)/2\n\n....\n\nI found that this method actually improves the RMSE score. I also tried this with more than one results, and interestingly it (almost everytime) gave improved results. \n\nSome may think this is just a trick, but I rather found it's statistically valid - since the scoring is based on RMSE, this method reduces variance within the results and reduce errors from outliers. So the result improves as much as the merging process reduces errors caused by variance itself. \n\nHowever, this raises me a question that 'is the result produced by this merging process could be seen as a **Good prediction result**?' And also this raised me a question of  'Is RMSE a good way to value preciseness of prediction result?'\n\nLow RMSE means in average, entities have low errors. It shows general tendencies of entire dataset. However, it doesn't count how much of the dataset got proper result, and how much haven't. \n\nSuppose that two predictions got same RMSE scores - One got entire rows incorrect, but to only little extent. The other got half of rows correct, but the other half incorrect to huge extent. Can we say these two prediction have identical validity in terms of correctness?\n\n\n(Anyway I believe this merging method could be a good way to tune your scores before submission for better scores!)\n\n  [1]: https://www.kaggle.com/moussaid/moyen-just-simple-moyen-of-my-predictions",
      "votes": 9
    },
    {
      "id": 339515,
      "postDate": "2018-06-07T04:32:44.230Z",
      "content": "<p>FWIW some simple tests find this to only make my score worse compared to simple averages of models.</p>",
      "rawMarkdown": "FWIW some simple tests find this to only make my score worse compared to simple averages of models.",
      "votes": 2,
      "replies": [
        {
          "id": 339521,
          "postDate": "2018-06-07T04:43:24.870Z",
          "content": "<p>I kept trying with several different cases, and it didn't improve for every cases. In my case trying with results from identical algorithm / identical codes didn't improve score. For me results from totally different process /algorithm jumped the score about 0.001</p>",
          "rawMarkdown": "I kept trying with several different cases, and it didn't improve for every cases. In my case trying with results from identical algorithm / identical codes didn't improve score. For me results from totally different process /algorithm jumped the score about 0.001"
        }
      ]
    },
    {
      "id": 339497,
      "postDate": "2018-06-07T03:46:50.910Z",
      "content": "<p>About your question with example of lots of small errors vs few large errors.\nRMSE from what I understand, takes care of that.\nMSE in the 2 cases will give same result but RMSE will give better score to lots of small errors vs few huge errors. \nI hope I'm not mistaken </p>",
      "rawMarkdown": "About your question with example of lots of small errors vs few large errors.\nRMSE from what I understand, takes care of that.\nMSE in the 2 cases will give same result but RMSE will give better score to lots of small errors vs few huge errors. \nI hope I'm not mistaken ",
      "votes": 2,
      "replies": [
        {
          "id": 339520,
          "postDate": "2018-06-07T04:41:20.800Z",
          "content": "<p>Oh thanks for the information! As a beginner I haven't thoroughly understood statistical concept of RMSE. That explains why that method gave better RMSE score!</p>",
          "rawMarkdown": "Oh thanks for the information! As a beginner I haven't thoroughly understood statistical concept of RMSE. That explains why that method gave better RMSE score!"
        }
      ]
    },
    {
      "id": 339473,
      "postDate": "2018-06-07T02:30:15.580Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 339518,
          "postDate": "2018-06-07T04:36:48.387Z",
          "content": "<p>Thanks for the remark! I agree with that assumption</p>",
          "rawMarkdown": "Thanks for the remark! I agree with that assumption"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 340501,
      "author_name": "Eric Bouteillon",
      "author_url": "",
      "post_date": "2018-06-09T11:26:13.963000",
      "content": "<p>Hi,</p>\n\n<p>From my perspective, you are just performing a weighted average on your predictions and you are lucky to find good weights.</p>\n\n<p>Here is why: let's say that you load predictions from two files x and y:</p>\n\n<p>d1 = x</p>\n\n<p>d2 = y</p>\n\n<p>Then I apply your algorithm:</p>\n\n<p>d1 = (d1+d2)/2 = (x+y)/2 = 1/2*x + 1/2*y</p>\n\n<p>d2 = (d1+d2)/2 = (1/2*x + 1/2*y + y)/2 = 1/4*x + 3/4*y</p>\n\n<p>d1 = (d1+d2)/2 = (1/2*x + 1/2*y + 1/4*x + 3/4*y)/2 = 3/8*x + 5/8*y</p>\n\n<p>d2 = (d1+d2)/2 = (3/8*x + 5/8*y + 1/4*x + 3/4*y)/2 = 5/16*x + 11/16*y</p>\n\n<p>....</p>\n\n<p>Which is a simple weighted average of x and y with some \"magic\" weights.\nMoreover, simply exchanging x and y will give you complete different results.</p>",
      "votes": 12,
      "replies": [
        {
          "id": 340739,
          "author_name": "Ethan Sukhyun Hong",
          "author_url": "",
          "post_date": "2018-06-10T07:21:22.563000",
          "content": "<p>I think that's a good point thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 339515,
      "author_name": "Peter Hurford",
      "author_url": "",
      "post_date": "2018-06-07T04:32:44.230000",
      "content": "<p>FWIW some simple tests find this to only make my score worse compared to simple averages of models.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 339521,
          "author_name": "Ethan Sukhyun Hong",
          "author_url": "",
          "post_date": "2018-06-07T04:43:24.870000",
          "content": "<p>I kept trying with several different cases, and it didn't improve for every cases. In my case trying with results from identical algorithm / identical codes didn't improve score. For me results from totally different process /algorithm jumped the score about 0.001</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 339497,
      "author_name": "AmirH",
      "author_url": "",
      "post_date": "2018-06-07T03:46:50.910000",
      "content": "<p>About your question with example of lots of small errors vs few large errors.\nRMSE from what I understand, takes care of that.\nMSE in the 2 cases will give same result but RMSE will give better score to lots of small errors vs few huge errors. \nI hope I'm not mistaken </p>",
      "votes": 2,
      "replies": [
        {
          "id": 339520,
          "author_name": "Ethan Sukhyun Hong",
          "author_url": "",
          "post_date": "2018-06-07T04:41:20.800000",
          "content": "<p>Oh thanks for the information! As a beginner I haven't thoroughly understood statistical concept of RMSE. That explains why that method gave better RMSE score!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 339473,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-07T02:30:15.580000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 339518,
          "author_name": "Ethan Sukhyun Hong",
          "author_url": "",
          "post_date": "2018-06-07T04:36:48.387000",
          "content": "<p>Thanks for the remark! I agree with that assumption</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "340501": "Hi,\n\nFrom my perspective, you are just performing a weighted average on your predictions and you are lucky to find good weights.\n\nHere is why: let's say that you load predictions from two files x and y:\n\nd1 = x\n\nd2 = y\n\nThen I apply your algorithm:\n\nd1 = (d1+d2)/2 = (x+y)/2 = 1/2*x + 1/2*y\n\nd2 = (d1+d2)/2 = (1/2*x + 1/2*y + y)/2 = 1/4*x + 3/4*y\n\nd1 = (d1+d2)/2 = (1/2*x + 1/2*y + 1/4*x + 3/4*y)/2 = 3/8*x + 5/8*y\n\nd2 = (d1+d2)/2 = (3/8*x + 5/8*y + 1/4*x + 3/4*y)/2 = 5/16*x + 11/16*y\n\n....\n\nWhich is a simple weighted average of x and y with some \"magic\" weights.\nMoreover, simply exchanging x and y will give you complete different results.",
    "339448": "I found an interesting [kernal :][1] \nThis one just try to find converging point between the two results - by repeat giving average of two results. For example, like this:\n\nimport d1, d2 (two results)\n\nd1 = (d1+d2)/2\n\nd2 = (d1+d2)/2\n\nd1 = (d1+d2)/2\n\nd2 = (d1+d2)/2\n\n....\n\nI found that this method actually improves the RMSE score. I also tried this with more than one results, and interestingly it (almost everytime) gave improved results. \n\nSome may think this is just a trick, but I rather found it's statistically valid - since the scoring is based on RMSE, this method reduces variance within the results and reduce errors from outliers. So the result improves as much as the merging process reduces errors caused by variance itself. \n\nHowever, this raises me a question that 'is the result produced by this merging process could be seen as a **Good prediction result**?' And also this raised me a question of  'Is RMSE a good way to value preciseness of prediction result?'\n\nLow RMSE means in average, entities have low errors. It shows general tendencies of entire dataset. However, it doesn't count how much of the dataset got proper result, and how much haven't. \n\nSuppose that two predictions got same RMSE scores - One got entire rows incorrect, but to only little extent. The other got half of rows correct, but the other half incorrect to huge extent. Can we say these two prediction have identical validity in terms of correctness?\n\n\n(Anyway I believe this merging method could be a good way to tune your scores before submission for better scores!)\n\n  [1]: https://www.kaggle.com/moussaid/moyen-just-simple-moyen-of-my-predictions",
    "339515": "FWIW some simple tests find this to only make my score worse compared to simple averages of models.",
    "339497": "About your question with example of lots of small errors vs few large errors.\nRMSE from what I understand, takes care of that.\nMSE in the 2 cases will give same result but RMSE will give better score to lots of small errors vs few huge errors. \nI hope I'm not mistaken ",
    "339473": ""
  }
}