{
  "id": 332286,
  "title": "The Random Kaggle estimator",
  "url": "/competitions/amex-default-prediction/discussion/332286",
  "author_name": "",
  "post_date": "2022-06-21T05:39:27.890393900Z",
  "votes": 43,
  "comment_count": 23,
  "views": 0,
  "content": "<p>Just a shower thought whilst running code; although ostensibly each of us are working alone to win this competition, we are in some way forming the \"Random Kaggle estimator\".</p>\n<p>Each team is a \"weak\" learner (not meant in any pejorative sense) akin to a decision tree estimator. However, collectively we are like the strong Random Forest leading to an almost optimal solution, using the Public Leaderboard as our objective function. We are very much guided by the work of others, either directly (for example forking and tweaking public notebooks) or indirectly (ideas as to what to explore further, or what others have already tried and do not seem to work, or even simply observing a sudden jump in the top LB score indicating that there is some <em>magic</em> to be found). The ideas and code do not even have to originate from this competition, but maybe are from a similar kaggle competition in the (perhaps distant) past, and some of the more recent Tabular Playground Series write-ups are astounding and truly inspiring!</p>\n<p>It is this 'human-in-the-loop' meta-estimator that can be found only on kaggle! </p>\n<p>Anyway, I think my script has just finished, so back to work…</p>\n<p>Good luck to all; the better you do, the better everyone does!</p>",
  "messages": [
    {
      "id": "1827527",
      "postDate": "06/21/2022 05:39:27",
      "content": "<p>Just a shower thought whilst running code; although ostensibly each of us are working alone to win this competition, we are in some way forming the \"Random Kaggle estimator\".</p>\n<p>Each team is a \"weak\" learner (not meant in any pejorative sense) akin to a decision tree estimator. However, collectively we are like the strong Random Forest leading to an almost optimal solution, using the Public Leaderboard as our objective function. We are very much guided by the work of others, either directly (for example forking and tweaking public notebooks) or indirectly (ideas as to what to explore further, or what others have already tried and do not seem to work, or even simply observing a sudden jump in the top LB score indicating that there is some <em>magic</em> to be found). The ideas and code do not even have to originate from this competition, but maybe are from a similar kaggle competition in the (perhaps distant) past, and some of the more recent Tabular Playground Series write-ups are astounding and truly inspiring!</p>\n<p>It is this 'human-in-the-loop' meta-estimator that can be found only on kaggle! </p>\n<p>Anyway, I think my script has just finished, so back to work…</p>\n<p>Good luck to all; the better you do, the better everyone does!</p>",
      "rawMarkdown": "Just a shower thought whilst running code; although ostensibly each of us are working alone to win this competition, we are in some way forming the \"Random Kaggle estimator\".\n\nEach team is a \"weak\" learner (not meant in any pejorative sense) akin to a decision tree estimator. However, collectively we are like the strong Random Forest leading to an almost optimal solution, using the Public Leaderboard as our objective function. We are very much guided by the work of others, either directly (for example forking and tweaking public notebooks) or indirectly (ideas as to what to explore further, or what others have already tried and do not seem to work, or even simply observing a sudden jump in the top LB score indicating that there is some *magic* to be found). The ideas and code do not even have to originate from this competition, but maybe are from a similar kaggle competition in the (perhaps distant) past, and some of the more recent Tabular Playground Series write-ups are astounding and truly inspiring!\n\nIt is this 'human-in-the-loop' meta-estimator that can be found only on kaggle! \n\nAnyway, I think my script has just finished, so back to work...\n\nGood luck to all; the better you do, the better everyone does!",
      "votes": null
    },
    {
      "id": "1827661",
      "postDate": "06/21/2022 07:58:36",
      "content": "<p>if all team can made their submission csv file public …. then it will be interesting</p>",
      "rawMarkdown": "if all team can made their submission csv file public .... then it will be interesting",
      "votes": null
    },
    {
      "id": "1827673",
      "postDate": "06/21/2022 08:09:53",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>\n<p>Amusing idea! One thing about stand-alone <code>submission.csv</code> files is that they are a bit like sausages; you don't know what went into making them. Indeed perhaps all closed-source blending notebooks should even come with a <em>caveat emptor</em>. For example, they must clearly say on the packet whether they have   <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328756\" target=\"_blank\"><code>B_29</code></a> added into the mix or not 😃</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @hengck23 \n\nAmusing idea! One thing about stand-alone `submission.csv` files is that they are a bit like sausages; you don't know what went into making them. Indeed perhaps all closed-source blending notebooks should even come with a *caveat emptor*. For example, they must clearly say on the packet whether they have   [`B_29`](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328756) added into the mix or not 😃\n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "1827784",
      "postDate": "06/21/2022 09:26:36",
      "content": "<p>We are all part of a big ensemble.</p>",
      "rawMarkdown": "We are all part of a big ensemble.",
      "votes": null
    },
    {
      "id": "1827816",
      "postDate": "06/21/2022 10:04:28",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a> I would be curious if there is some gain in averaging lots of score or if it is somehow capped. I would be also curious of innovative weighting scheme that could be used. I know that some market finance techniques (Hierarchical Risk Parity) was proposed.</p>",
      "rawMarkdown": "Hi @carlmcbrideellis I would be curious if there is some gain in averaging lots of score or if it is somehow capped. I would be also curious of innovative weighting scheme that could be used. I know that some market finance techniques (Hierarchical Risk Parity) was proposed.",
      "votes": null
    },
    {
      "id": "1827819",
      "postDate": "06/21/2022 10:09:33",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/lucasmorin\" target=\"_blank\">@lucasmorin</a> </p>\n<p>I am completely unfamiliar with \"Hierarchical Risk Parity\", do you have any links?</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @lucasmorin \n\nI am completely unfamiliar with \"Hierarchical Risk Parity\", do you have any links?\n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "1827849",
      "postDate": "06/21/2022 10:31:47",
      "content": "<p>well at least <code>B_29</code> is at the bottom of the list in terms how useful it is: )</p>",
      "rawMarkdown": "well at least `B_29` is at the bottom of the list in terms how useful it is: )",
      "votes": null
    },
    {
      "id": "1827856",
      "postDate": "06/21/2022 10:38:39",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> </p>\n<p>Only for the training data and the Public chunk; ¿what happens when it 'kicks-in' on the Private data is the question?</p>\n<p>That said, if the estimators they use employ just a smidgen of regularization (<em>i.e.</em> L2, or perhaps better still L1) they probably  create a model that diminishes (L2) or disregards (L1) the feature  <code>B_29</code> anyway…</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @raddar \n\nOnly for the training data and the Public chunk; ¿what happens when it 'kicks-in' on the Private data is the question?\n\nThat said, if the estimators they use employ just a smidgen of regularization (*i.e.* L2, or perhaps better still L1) they probably  create a model that diminishes (L2) or disregards (L1) the feature  `B_29` anyway...\n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "1827938",
      "postDate": "06/21/2022 12:06:16",
      "content": "<p>HRP Paper: <a href=\"https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2708678\" target=\"_blank\">https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2708678</a><br>\nHRP Slides: <a href=\"https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2713516\" target=\"_blank\">https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2713516</a></p>\n<p>I've seen the idea of applying portfolio selection approach to model ensembling as the problem are relatively similar (can't find a good ressource right now).</p>",
      "rawMarkdown": "HRP Paper: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2708678\nHRP Slides: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2713516\n\nI've seen the idea of applying portfolio selection approach to model ensembling as the problem are relatively similar (can't find a good ressource right now).",
      "votes": null
    },
    {
      "id": "1828000",
      "postDate": "06/21/2022 12:50:08",
      "content": "<h3>Kaggle Random Walk</h3>\n<h5>Regularized Stochastic Leaderboard Descent</h5>\n<hr>\n<h3>Abstract</h3>\n<p>We demonstrate the real work behind winning any Kaggle competition is actually accomplished through a human-computer collaboration. Teams are treated as individual weak learners who compete against one another by deriving optimized estimation solutions on the public leaderboard. Mutually and unconsciously, competitors in turn capitalize on the strengths and mitigate the weaknesses of other competitors by re-reporting their findings, when necessary, and ultimately developing their own solutions. By assembling the individual solutions together, we obtain an ideal solution, which is almost optimal. This \"Random Kaggle estimator\" enables us to demonstrate how humans, as part of an ML system, significantly contribute to its overall ability to solve complex problems. The more competitors, the more distinct differences there are between them (a min. two requisite), and the greater the leaderboard fluctuations, the faster and better the meta-estimation solution can be developed.</p>\n<p>Furthermore, we define a measure of the degree to which random actions by humans are required to accomplish a task relative to the task's complexity - \"Stochastic Leaderboard Descent\". We will demonstrate that the more complex a problem-solving task requires a greater number of independent teams working in unison to accomplish the ML task, with increasing importance placed on each team as the degree of stochasticity increases. We also show that although some of the ideas and code do not originate from the same competition (or are from the same competition at a different time), there are still ideas that we can not attribute to being influenced by any specific historical source.</p>\n<p>Suggesting that such ideas result from pure creativity and are driven by intuition, a human strength that extends far beyond ML tasks.</p>\n<h4>Introduction</h4>\n<p>In the world of Kaggle, very often only the top spot on the leaderboard matters. On the other hand, once the leaderboard is released, it becomes a valuable tool for researchers to learn from and obtain new insights, especially in competitions with a sufficient number of competitors.</p>\n<p>In this paper, we show that there is much more to leaderboards than the winner's name. Using a recent Kaggle competition as an example, we illustrate that competitors are able to adapt to one another and progressively improve their leaderboard score by benefiting from the cumulative effect of multiple contributions. We formulate a \"Random Kaggle estimator\" which, although simple, is able to reproduce the leaderboard evolution curve and we derive some insights from this model.</p>\n<p>We then demonstrate how the leaderboard evolution curve itself can be used to define an optimal estimation model, which we call the \"stochastic leaderboard descent\".</p>\n<h3>Preliminaries</h3>\n<h4>Generalized Blending Theory</h4>\n<p><strong>Definition 1: Solution Blending</strong></p>\n<blockquote>\n  <p>The process of assembling the predictions of multiple solutions in order to produce a new and improved estimate of the target value.</p>\n</blockquote>\n<p>Therefore, it follows from Definition 1 that a blended solution is the combination of multiple solutions that can be assembled to produce a more accurate estimate of the target value. As the blended solution is a combination of multiple solutions, it is a collaborative ensemble of solutions. Therefore, we can define a blended solution as follows:</p>\n<p>$$\\beta(x_i) = \\sum_j \\alpha_j(x_i) x_j(x_i)$$</p>\n<p>Where <em>β</em> is the blended solution, <em>𝛼</em> is the solution weight and <em>x_j(x_i)</em> is the prediction for the training instance <em>x_i</em>.</p>\n<p><strong>Ideas Blending</strong></p>\n<p>Before we define we must first define <em>Idea</em>:</p>\n<p>Data scientists come up with a metric and then optimize that metric.</p>\n<p>As we can see the information of the target outcome is linked to both the solution and the methods used to optimize the solution. We can represent a solution as follows:</p>\n<p>$$x_i(z) = \\Phi_i(z)\\phi_i(z)$$</p>\n<p>Where <em>z</em> is the data, <em>Φ</em> is the encoder (Optimization method) and 𝜙 is the decoder (model), the feature-based solution.</p>\n<p>Lastly, we are able to formulate the idea of data scientist as:</p>\n<p>$$y_i(z) = M(x_i(z))$$</p>\n<p>Where <em>y_i(z)</em> is the data science outcome and <em>M</em> is the information linking the respective solution to the metric.</p>\n<p><strong>Definition 2: Ideas Blending</strong></p>\n<blockquote>\n  <p>The process of assembling the predictions of multiple ideas in order to produce a new and improved estimate of the target value.</p>\n</blockquote>\n<p>From Definition 1, we can define <em>Ideas Blending</em> as:</p>\n<p>$$y_i = \\arg \\min_{\\beta(z)} \\left[\\sum_j \\alpha_j y_j(z)\\right]$$</p>\n<p>Since the scientists try to blend ideas in order to minimize his leaderboard position.</p>\n<p><strong>Definition 3: Stochastic Leaderboard Convergence (SLC)</strong></p>\n<blockquote>\n  <p>A Stochastic Leaderboard Convergence is a process of obtaining an improved estimate of the target value by re-estimating the target value after a new public leaderboard score has been reported.</p>\n</blockquote>\n<p>It follows from Definition 1 that a Stochastic Leaderboard Convergence (SLC) is a process of obtaining an improved estimate of the target value by re-estimating the target value after a new public leaderboard score has been reported. The idea behind SLC is that the more estimates of the target value are reported on the public leaderboard, the better the estimate of the target value will be.</p>\n<p>We can therefore formulate the Stochastic Leaderboard Convergence (SLC) as follows:</p>\n<p>$$\\hat{y} = \\arg \\min_{\\beta(x_i)} \\left[\\sum_i \\left[\\sum_j \\alpha_j(x_i) x_j(x_i) - y_i\\right]^2\\right]$$</p>\n<p><strong>Definition 4: Public Leaderboard Density</strong></p>\n<blockquote>\n  <p>The density of the number of public leaderboard scores reported.</p>\n</blockquote>\n<p>The Public Leaderboard Density (PLD) is the number of reported public leaderboard scores per unit area (competition timeframe). We can formulate the PLD as follows:</p>\n<p>$$PLD = \\frac{N}{T}$$</p>\n<p>Where <em>N</em> is the number of public leaderboard scores reported, and <em>T</em> is the competition timeframe.</p>\n<p><strong>Definition 4: Stochastic Leaderboard Descent (SLD)</strong></p>\n<blockquote>\n  <p>A Stochastic Leaderboard Descent is a process of obtaining an improved estimate of the target value by taking the current public leaderboard score as the starting point and applying stochastic gradient descent to obtain a better estimate of the target value.</p>\n</blockquote>\n<h4>Leaderboard Monte Carlo Estimation</h4>\n<p><strong>Bayesian Blending</strong></p>\n<p>Bayesian Blending is a process of obtaining an improved estimate of the target value by taking the current top public solution as the starting point and applying stochastic notebook descent to obtain a better estimate of the target value.<br>\nSince there are: </p>\n<ul>\n<li><strong>Solutions</strong> released as notebooks</li>\n<li><strong>Ideas</strong> released on the discussions forums</li>\n</ul>\n<p>We propse to define the blended solution as a linear combination of both a <code>solution term</code> and an <code>idea term</code>:</p>\n<p>$$\\omega(x_i) = \\sum_j \\alpha_j n_j(z) + \\sum_k \\beta_k M(x_k(z))$$</p>\n<p>Where <em>n_j</em> is the notebook's code used (function of the data <em>z</em>) and <em>M(x_k(z))</em> again is the information extracted from an idea <em>x_i</em> (function of the competition's data). The weights <em>𝛼</em> and <em>𝛽</em> are the probabilistic weight of the notebooks and ideas.</p>\n<p>We therefore can define the direction of the kaggle bayesian blending process as the gradient of the notebook-idea-blended solution.</p>\n<p>$$\\nabla_k(z) = \\sum_j \\alpha_j(z) \\nabla_k n_j(z) + \\sum_j \\beta_j(z) \\nabla_k M(x_j(z))$$</p>\n<p>Therefore, the notebook idea blended solution can be formulated as a kernel function:</p>\n<p>$$K_j(z,z') = \\nabla_k(z) x_j(z')$$</p>\n<p>To optimize the kernel function, we can introduce a kernel matrix of all the kernels and their respective gradients for each notebook-idea-blended solution.</p>\n<p>$$\\hat{K}_{jk}(z,z') = \\nabla_k(z) x_j(z')$$</p>\n<h4>Past solution influence</h4>\n<p>Historical influence can sometime be attributed to some public submissions. Also with the recent social activity made popular by covid lockdown elimination this type of influence becomes more and more of a factor. Thus, from second law of medaldynamics we can assume that just like entropy, ideas flow from low medal density areas to high medal density areas, so we can define the past idea-solution as:</p>\n<p>$$n_j(z) = \\sum_i \\nabla_{i}(z)$$</p>\n<p>Where <em>𝛼</em> is the coefficient for idea divergence <em>(The second law of medaldynamics)</em> and <em>𝛽</em> is the past idea-solution.</p>\n<p>We can therefore define the <strong>emergence of creativity</strong> as the process of extracting the solution that is orthogonal to past influence.</p>",
      "rawMarkdown": "### Kaggle Random Walk\n##### Regularized Stochastic Leaderboard Descent\n_____\n\n### Abstract\nWe demonstrate the real work behind winning any Kaggle competition is actually accomplished through a human-computer collaboration. Teams are treated as individual weak learners who compete against one another by deriving optimized estimation solutions on the public leaderboard. Mutually and unconsciously, competitors in turn capitalize on the strengths and mitigate the weaknesses of other competitors by re-reporting their findings, when necessary, and ultimately developing their own solutions. By assembling the individual solutions together, we obtain an ideal solution, which is almost optimal. This \"Random Kaggle estimator\" enables us to demonstrate how humans, as part of an ML system, significantly contribute to its overall ability to solve complex problems. The more competitors, the more distinct differences there are between them (a min. two requisite), and the greater the leaderboard fluctuations, the faster and better the meta-estimation solution can be developed.\n\nFurthermore, we define a measure of the degree to which random actions by humans are required to accomplish a task relative to the task's complexity - \"Stochastic Leaderboard Descent\". We will demonstrate that the more complex a problem-solving task requires a greater number of independent teams working in unison to accomplish the ML task, with increasing importance placed on each team as the degree of stochasticity increases. We also show that although some of the ideas and code do not originate from the same competition (or are from the same competition at a different time), there are still ideas that we can not attribute to being influenced by any specific historical source.\n\nSuggesting that such ideas result from pure creativity and are driven by intuition, a human strength that extends far beyond ML tasks.\n\n#### Introduction\n\nIn the world of Kaggle, very often only the top spot on the leaderboard matters. On the other hand, once the leaderboard is released, it becomes a valuable tool for researchers to learn from and obtain new insights, especially in competitions with a sufficient number of competitors.\n\nIn this paper, we show that there is much more to leaderboards than the winner's name. Using a recent Kaggle competition as an example, we illustrate that competitors are able to adapt to one another and progressively improve their leaderboard score by benefiting from the cumulative effect of multiple contributions. We formulate a \"Random Kaggle estimator\" which, although simple, is able to reproduce the leaderboard evolution curve and we derive some insights from this model.\n\nWe then demonstrate how the leaderboard evolution curve itself can be used to define an optimal estimation model, which we call the \"stochastic leaderboard descent\".\n\n### Preliminaries\n\n#### Generalized Blending Theory\n\n**Definition 1: Solution Blending**\n> The process of assembling the predictions of multiple solutions in order to produce a new and improved estimate of the target value.\n\nTherefore, it follows from Definition 1 that a blended solution is the combination of multiple solutions that can be assembled to produce a more accurate estimate of the target value. As the blended solution is a combination of multiple solutions, it is a collaborative ensemble of solutions. Therefore, we can define a blended solution as follows:\n\n$$\\beta(x_i) = \\sum_j \\alpha_j(x_i) x_j(x_i)$$\n\nWhere *β* is the blended solution, *𝛼* is the solution weight and *x_j(x_i)* is the prediction for the training instance *x_i*.\n\n**Ideas Blending**\n\nBefore we define we must first define *Idea*:\n\nData scientists come up with a metric and then optimize that metric.\n\nAs we can see the information of the target outcome is linked to both the solution and the methods used to optimize the solution. We can represent a solution as follows:\n\n$$x_i(z) = \\Phi_i(z)\\phi_i(z)$$\n\nWhere *z* is the data, *Φ* is the encoder (Optimization method) and 𝜙 is the decoder (model), the feature-based solution.\n\nLastly, we are able to formulate the idea of data scientist as:\n\n$$y_i(z) = M(x_i(z))$$\n\nWhere *y_i(z)* is the data science outcome and *M* is the information linking the respective solution to the metric.\n\n**Definition 2: Ideas Blending**\n> The process of assembling the predictions of multiple ideas in order to produce a new and improved estimate of the target value.\n\nFrom Definition 1, we can define *Ideas Blending* as:\n\n$$y_i = \\arg \\min_{\\beta(z)} \\left[\\sum_j \\alpha_j y_j(z)\\right]$$\n\nSince the scientists try to blend ideas in order to minimize his leaderboard position.\n\n**Definition 3: Stochastic Leaderboard Convergence (SLC)**\n> A Stochastic Leaderboard Convergence is a process of obtaining an improved estimate of the target value by re-estimating the target value after a new public leaderboard score has been reported.\n\nIt follows from Definition 1 that a Stochastic Leaderboard Convergence (SLC) is a process of obtaining an improved estimate of the target value by re-estimating the target value after a new public leaderboard score has been reported. The idea behind SLC is that the more estimates of the target value are reported on the public leaderboard, the better the estimate of the target value will be.\n\nWe can therefore formulate the Stochastic Leaderboard Convergence (SLC) as follows:\n\n$$\\hat{y} = \\arg \\min_{\\beta(x_i)} \\left[\\sum_i \\left[\\sum_j \\alpha_j(x_i) x_j(x_i) - y_i\\right]^2\\right]$$\n\n**Definition 4: Public Leaderboard Density**\n> The density of the number of public leaderboard scores reported.\n\nThe Public Leaderboard Density (PLD) is the number of reported public leaderboard scores per unit area (competition timeframe). We can formulate the PLD as follows:\n\n$$PLD = \\frac{N}{T}$$\n\nWhere *N* is the number of public leaderboard scores reported, and *T* is the competition timeframe.\n\n**Definition 4: Stochastic Leaderboard Descent (SLD)**\n> A Stochastic Leaderboard Descent is a process of obtaining an improved estimate of the target value by taking the current public leaderboard score as the starting point and applying stochastic gradient descent to obtain a better estimate of the target value.\n\n#### Leaderboard Monte Carlo Estimation\n\n**Bayesian Blending**\n\nBayesian Blending is a process of obtaining an improved estimate of the target value by taking the current top public solution as the starting point and applying stochastic notebook descent to obtain a better estimate of the target value.\nSince there are: \n\n- **Solutions** released as notebooks\n- **Ideas** released on the discussions forums\n\nWe propse to define the blended solution as a linear combination of both a `solution term` and an `idea term`:\n\n$$\\omega(x_i) = \\sum_j \\alpha_j n_j(z) + \\sum_k \\beta_k M(x_k(z))$$\n\nWhere *n_j* is the notebook's code used (function of the data *z*) and *M(x_k(z))* again is the information extracted from an idea *x_i* (function of the competition's data). The weights *𝛼* and *𝛽* are the probabilistic weight of the notebooks and ideas.\n\nWe therefore can define the direction of the kaggle bayesian blending process as the gradient of the notebook-idea-blended solution.\n\n$$\\nabla_k(z) = \\sum_j \\alpha_j(z) \\nabla_k n_j(z) + \\sum_j \\beta_j(z) \\nabla_k M(x_j(z))$$\n\nTherefore, the notebook idea blended solution can be formulated as a kernel function:\n\n$$K_j(z,z') = \\nabla_k(z) x_j(z')$$\n\nTo optimize the kernel function, we can introduce a kernel matrix of all the kernels and their respective gradients for each notebook-idea-blended solution.\n\n$$\\hat{K}_{jk}(z,z') = \\nabla_k(z) x_j(z')$$\n\n#### Past solution influence\n\nHistorical influence can sometime be attributed to some public submissions. Also with the recent social activity made popular by covid lockdown elimination this type of influence becomes more and more of a factor. Thus, from second law of medaldynamics we can assume that just like entropy, ideas flow from low medal density areas to high medal density areas, so we can define the past idea-solution as:\n\n$$n_j(z) = \\sum_i \\nabla_{i}(z)$$\n\nWhere *𝛼* is the coefficient for idea divergence *(The second law of medaldynamics)* and *𝛽* is the past idea-solution.\n\nWe can therefore define the **emergence of creativity** as the process of extracting the solution that is orthogonal to past influence.",
      "votes": null
    },
    {
      "id": "1828038",
      "postDate": "06/21/2022 13:03:12",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> </p>\n<p>Er, wow! I guess that is why you are called \"The Devastator\"… 😃</p>\n<p>That said, somewhat more than your \"<em>Regularized Stochastic Leaderboard Descent</em>\"(!) I am sort of reminded of the <a href=\"https://en.wikipedia.org/wiki/Genetic_algorithm\" target=\"_blank\">genetic algorithms</a> that were very much the fashion some 20-30 years ago….</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @thedevastator \n\nEr, wow! I guess that is why you are called \"The Devastator\"... 😃\n\nThat said, somewhat more than your \"*Regularized Stochastic Leaderboard Descent*\"(!) I am sort of reminded of the [genetic algorithms](https://en.wikipedia.org/wiki/Genetic_algorithm) that were very much the fashion some 20-30 years ago....\n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "1828351",
      "postDate": "06/21/2022 16:20:54",
      "content": "<p>Kagglers…Assemble.</p>",
      "rawMarkdown": "Kagglers...Assemble.",
      "votes": null
    },
    {
      "id": "1828393",
      "postDate": "06/21/2022 16:45:49",
      "content": "<p>You should get likes simply for attempting to write this!!🙌</p>",
      "rawMarkdown": "You should get likes simply for attempting to write this!!🙌",
      "votes": null
    },
    {
      "id": "1828405",
      "postDate": "06/21/2022 16:58:26",
      "content": "<p>IIRC, in the past, organizers have created aggregate models from submissions after the end of some competitions. I don’t remember which competititions but it might have been some of the basketball or american football ones.</p>",
      "rawMarkdown": "IIRC, in the past, organizers have created aggregate models from submissions after the end of some competitions. I don’t remember which competititions but it might have been some of the basketball or american football ones.",
      "votes": null
    },
    {
      "id": "1828408",
      "postDate": "06/21/2022 17:02:09",
      "content": "<p>Thank you, Thank you 😁</p>",
      "rawMarkdown": "Thank you, Thank you 😁",
      "votes": null
    },
    {
      "id": "1828419",
      "postDate": "06/21/2022 17:11:18",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/christoffer\" target=\"_blank\">@christoffer</a> </p>\n<p>That is interesting. Although from an organizers point of view it is worth distinguishing between ensembling a diverse set of models using the provided code (as per the <a href=\"https://web.archive.org/web/20160304031055/http://mlwave.com/kaggle-ensembling-guide/\" target=\"_blank\">\"Kaggle Ensembling Guide\"</a>) which can (could? <a href=\"https://www.wired.com/2012/04/netflix-prize-costs/\" target=\"_blank\">\"<em>Netflix Never Used Its $1 Million Algorithm Due To Engineering Costs</em>\"</a>) be beneficial to the organizers, as opposed to the blending of raw <code>submission.csv</code> files together, which may also do well on a kaggle  competition, but on the other hand provide little (basically nothing) of use to the organizers, as they are solutions specific to the dataset provided. (PS: I am a little surprised that this is not a \"Code\" competition).</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @christoffer \n\nThat is interesting. Although from an organizers point of view it is worth distinguishing between ensembling a diverse set of models using the provided code (as per the [\"Kaggle Ensembling Guide\"](https://web.archive.org/web/20160304031055/http://mlwave.com/kaggle-ensembling-guide/)) which can (could? [\"*Netflix Never Used Its $1 Million Algorithm Due To Engineering Costs*\"](https://www.wired.com/2012/04/netflix-prize-costs/)) be beneficial to the organizers, as opposed to the blending of raw `submission.csv` files together, which may also do well on a kaggle  competition, but on the other hand provide little (basically nothing) of use to the organizers, as they are solutions specific to the dataset provided. (PS: I am a little surprised that this is not a \"Code\" competition).\n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "1829214",
      "postDate": "06/22/2022 12:55:47",
      "content": "<p>Thanks for that link.  I've been learning much more about ensembles and missed that resource.</p>",
      "rawMarkdown": "Thanks for that link.  I've been learning much more about ensembles and missed that resource.",
      "votes": null
    },
    {
      "id": "1829757",
      "postDate": "06/22/2022 22:58:44",
      "content": "<p>Beautifully written. I scout the site for technical insights but, occasionally, I stumble on such meta posts which make the process not just useful, but also pleasurable.</p>",
      "rawMarkdown": "Beautifully written. I scout the site for technical insights but, occasionally, I stumble on such meta posts which make the process not just useful, but also pleasurable.",
      "votes": null
    },
    {
      "id": "1829985",
      "postDate": "06/23/2022 04:54:26",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/edwardakalarrywelch\" target=\"_blank\">@edwardakalarrywelch</a> </p>\n<p>Motivated by your comment I have just posted the link as a Topic. As this is the first big \"Featured\" non-forecasting tabular competition in quite a while, and could hit 4-5k participants, and there will be quite a few people who have not seen \"kaggle ensembling\" in the text books. Furthermore, the site MLwave went down about a year ago, but thankfully the content was rescued by the Wayback Machine. </p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @edwardakalarrywelch \n\nMotivated by your comment I have just posted the link as a Topic. As this is the first big \"Featured\" non-forecasting tabular competition in quite a while, and could hit 4-5k participants, and there will be quite a few people who have not seen \"kaggle ensembling\" in the text books. Furthermore, the site MLwave went down about a year ago, but thankfully the content was rescued by the Wayback Machine. \n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "1830437",
      "postDate": "06/23/2022 12:49:08",
      "content": "<p>People at AMEX were probably like: <br>\n— Hey we need an insane model for this business problem…<br>\n— Lets just fit a Kaggle Random Forest.</p>",
      "rawMarkdown": "People at AMEX were probably like: \n— Hey we need an insane model for this business problem...\n— Lets just fit a Kaggle Random Forest.",
      "votes": null
    },
    {
      "id": "1830440",
      "postDate": "06/23/2022 12:53:17",
      "content": "<p>Hola <a href=\"https://www.kaggle.com/heyspaceturtle\" target=\"_blank\">@heyspaceturtle</a> </p>\n<p>Maybe, but it takes 3 months to run, and doesn't come cheap with $100,000 for prize money and an unknown amount in compute + storage…</p>\n<p>That said, credit default could cost somebody like American Express $M each year, so even a slightly better model could be worth the outlay.</p>\n<p>Un saludo muy codrial,<br>\ncarl</p>",
      "rawMarkdown": "Hola @heyspaceturtle \n\nMaybe, but it takes 3 months to run, and doesn't come cheap with $100,000 for prize money and an unknown amount in compute + storage...\n\nThat said, credit default could cost somebody like American Express $M each year, so even a slightly better model could be worth the outlay.\n\nUn saludo muy codrial,\ncarl",
      "votes": null
    },
    {
      "id": "1830464",
      "postDate": "06/23/2022 13:06:50",
      "content": "<p>True… but they still get the chance to find talented data scientists for their team :) </p>\n<p>Saludos! </p>",
      "rawMarkdown": "True... but they still get the chance to find talented data scientists for their team :) \n\nSaludos!",
      "votes": null
    },
    {
      "id": "1832792",
      "postDate": "06/25/2022 11:54:52",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a>, upvoted :)</p>",
      "rawMarkdown": "Great work @thedevastator, upvoted :)",
      "votes": null
    },
    {
      "id": "1836195",
      "postDate": "06/28/2022 12:50:06",
      "content": "<p>nice one…!</p>",
      "rawMarkdown": "nice one...!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1827661,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/21/2022 07:58:36",
      "content": "<p>if all team can made their submission csv file public …. then it will be interesting</p>",
      "votes": null,
      "replies": [
        {
          "id": 1827673,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "06/21/2022 08:09:53",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>\n<p>Amusing idea! One thing about stand-alone <code>submission.csv</code> files is that they are a bit like sausages; you don't know what went into making them. Indeed perhaps all closed-source blending notebooks should even come with a <em>caveat emptor</em>. For example, they must clearly say on the packet whether they have   <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328756\" target=\"_blank\"><code>B_29</code></a> added into the mix or not 😃</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1827849,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "06/21/2022 10:31:47",
          "content": "<p>well at least <code>B_29</code> is at the bottom of the list in terms how useful it is: )</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1827856,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "06/21/2022 10:38:39",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> </p>\n<p>Only for the training data and the Public chunk; ¿what happens when it 'kicks-in' on the Private data is the question?</p>\n<p>That said, if the estimators they use employ just a smidgen of regularization (<em>i.e.</em> L2, or perhaps better still L1) they probably  create a model that diminishes (L2) or disregards (L1) the feature  <code>B_29</code> anyway…</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1827784,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "06/21/2022 09:26:36",
      "content": "<p>We are all part of a big ensemble.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1827816,
      "author_name": "lucasmorin",
      "author_url": "",
      "post_date": "06/21/2022 10:04:28",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a> I would be curious if there is some gain in averaging lots of score or if it is somehow capped. I would be also curious of innovative weighting scheme that could be used. I know that some market finance techniques (Hierarchical Risk Parity) was proposed.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1827819,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "06/21/2022 10:09:33",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/lucasmorin\" target=\"_blank\">@lucasmorin</a> </p>\n<p>I am completely unfamiliar with \"Hierarchical Risk Parity\", do you have any links?</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1827938,
          "author_name": "lucasmorin",
          "author_url": "",
          "post_date": "06/21/2022 12:06:16",
          "content": "<p>HRP Paper: <a href=\"https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2708678\" target=\"_blank\">https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2708678</a><br>\nHRP Slides: <a href=\"https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2713516\" target=\"_blank\">https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2713516</a></p>\n<p>I've seen the idea of applying portfolio selection approach to model ensembling as the problem are relatively similar (can't find a good ressource right now).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1828000,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "06/21/2022 12:50:08",
      "content": "<h3>Kaggle Random Walk</h3>\n<h5>Regularized Stochastic Leaderboard Descent</h5>\n<hr>\n<h3>Abstract</h3>\n<p>We demonstrate the real work behind winning any Kaggle competition is actually accomplished through a human-computer collaboration. Teams are treated as individual weak learners who compete against one another by deriving optimized estimation solutions on the public leaderboard. Mutually and unconsciously, competitors in turn capitalize on the strengths and mitigate the weaknesses of other competitors by re-reporting their findings, when necessary, and ultimately developing their own solutions. By assembling the individual solutions together, we obtain an ideal solution, which is almost optimal. This \"Random Kaggle estimator\" enables us to demonstrate how humans, as part of an ML system, significantly contribute to its overall ability to solve complex problems. The more competitors, the more distinct differences there are between them (a min. two requisite), and the greater the leaderboard fluctuations, the faster and better the meta-estimation solution can be developed.</p>\n<p>Furthermore, we define a measure of the degree to which random actions by humans are required to accomplish a task relative to the task's complexity - \"Stochastic Leaderboard Descent\". We will demonstrate that the more complex a problem-solving task requires a greater number of independent teams working in unison to accomplish the ML task, with increasing importance placed on each team as the degree of stochasticity increases. We also show that although some of the ideas and code do not originate from the same competition (or are from the same competition at a different time), there are still ideas that we can not attribute to being influenced by any specific historical source.</p>\n<p>Suggesting that such ideas result from pure creativity and are driven by intuition, a human strength that extends far beyond ML tasks.</p>\n<h4>Introduction</h4>\n<p>In the world of Kaggle, very often only the top spot on the leaderboard matters. On the other hand, once the leaderboard is released, it becomes a valuable tool for researchers to learn from and obtain new insights, especially in competitions with a sufficient number of competitors.</p>\n<p>In this paper, we show that there is much more to leaderboards than the winner's name. Using a recent Kaggle competition as an example, we illustrate that competitors are able to adapt to one another and progressively improve their leaderboard score by benefiting from the cumulative effect of multiple contributions. We formulate a \"Random Kaggle estimator\" which, although simple, is able to reproduce the leaderboard evolution curve and we derive some insights from this model.</p>\n<p>We then demonstrate how the leaderboard evolution curve itself can be used to define an optimal estimation model, which we call the \"stochastic leaderboard descent\".</p>\n<h3>Preliminaries</h3>\n<h4>Generalized Blending Theory</h4>\n<p><strong>Definition 1: Solution Blending</strong></p>\n<blockquote>\n  <p>The process of assembling the predictions of multiple solutions in order to produce a new and improved estimate of the target value.</p>\n</blockquote>\n<p>Therefore, it follows from Definition 1 that a blended solution is the combination of multiple solutions that can be assembled to produce a more accurate estimate of the target value. As the blended solution is a combination of multiple solutions, it is a collaborative ensemble of solutions. Therefore, we can define a blended solution as follows:</p>\n<p>$$\\beta(x_i) = \\sum_j \\alpha_j(x_i) x_j(x_i)$$</p>\n<p>Where <em>β</em> is the blended solution, <em>𝛼</em> is the solution weight and <em>x_j(x_i)</em> is the prediction for the training instance <em>x_i</em>.</p>\n<p><strong>Ideas Blending</strong></p>\n<p>Before we define we must first define <em>Idea</em>:</p>\n<p>Data scientists come up with a metric and then optimize that metric.</p>\n<p>As we can see the information of the target outcome is linked to both the solution and the methods used to optimize the solution. We can represent a solution as follows:</p>\n<p>$$x_i(z) = \\Phi_i(z)\\phi_i(z)$$</p>\n<p>Where <em>z</em> is the data, <em>Φ</em> is the encoder (Optimization method) and 𝜙 is the decoder (model), the feature-based solution.</p>\n<p>Lastly, we are able to formulate the idea of data scientist as:</p>\n<p>$$y_i(z) = M(x_i(z))$$</p>\n<p>Where <em>y_i(z)</em> is the data science outcome and <em>M</em> is the information linking the respective solution to the metric.</p>\n<p><strong>Definition 2: Ideas Blending</strong></p>\n<blockquote>\n  <p>The process of assembling the predictions of multiple ideas in order to produce a new and improved estimate of the target value.</p>\n</blockquote>\n<p>From Definition 1, we can define <em>Ideas Blending</em> as:</p>\n<p>$$y_i = \\arg \\min_{\\beta(z)} \\left[\\sum_j \\alpha_j y_j(z)\\right]$$</p>\n<p>Since the scientists try to blend ideas in order to minimize his leaderboard position.</p>\n<p><strong>Definition 3: Stochastic Leaderboard Convergence (SLC)</strong></p>\n<blockquote>\n  <p>A Stochastic Leaderboard Convergence is a process of obtaining an improved estimate of the target value by re-estimating the target value after a new public leaderboard score has been reported.</p>\n</blockquote>\n<p>It follows from Definition 1 that a Stochastic Leaderboard Convergence (SLC) is a process of obtaining an improved estimate of the target value by re-estimating the target value after a new public leaderboard score has been reported. The idea behind SLC is that the more estimates of the target value are reported on the public leaderboard, the better the estimate of the target value will be.</p>\n<p>We can therefore formulate the Stochastic Leaderboard Convergence (SLC) as follows:</p>\n<p>$$\\hat{y} = \\arg \\min_{\\beta(x_i)} \\left[\\sum_i \\left[\\sum_j \\alpha_j(x_i) x_j(x_i) - y_i\\right]^2\\right]$$</p>\n<p><strong>Definition 4: Public Leaderboard Density</strong></p>\n<blockquote>\n  <p>The density of the number of public leaderboard scores reported.</p>\n</blockquote>\n<p>The Public Leaderboard Density (PLD) is the number of reported public leaderboard scores per unit area (competition timeframe). We can formulate the PLD as follows:</p>\n<p>$$PLD = \\frac{N}{T}$$</p>\n<p>Where <em>N</em> is the number of public leaderboard scores reported, and <em>T</em> is the competition timeframe.</p>\n<p><strong>Definition 4: Stochastic Leaderboard Descent (SLD)</strong></p>\n<blockquote>\n  <p>A Stochastic Leaderboard Descent is a process of obtaining an improved estimate of the target value by taking the current public leaderboard score as the starting point and applying stochastic gradient descent to obtain a better estimate of the target value.</p>\n</blockquote>\n<h4>Leaderboard Monte Carlo Estimation</h4>\n<p><strong>Bayesian Blending</strong></p>\n<p>Bayesian Blending is a process of obtaining an improved estimate of the target value by taking the current top public solution as the starting point and applying stochastic notebook descent to obtain a better estimate of the target value.<br>\nSince there are: </p>\n<ul>\n<li><strong>Solutions</strong> released as notebooks</li>\n<li><strong>Ideas</strong> released on the discussions forums</li>\n</ul>\n<p>We propse to define the blended solution as a linear combination of both a <code>solution term</code> and an <code>idea term</code>:</p>\n<p>$$\\omega(x_i) = \\sum_j \\alpha_j n_j(z) + \\sum_k \\beta_k M(x_k(z))$$</p>\n<p>Where <em>n_j</em> is the notebook's code used (function of the data <em>z</em>) and <em>M(x_k(z))</em> again is the information extracted from an idea <em>x_i</em> (function of the competition's data). The weights <em>𝛼</em> and <em>𝛽</em> are the probabilistic weight of the notebooks and ideas.</p>\n<p>We therefore can define the direction of the kaggle bayesian blending process as the gradient of the notebook-idea-blended solution.</p>\n<p>$$\\nabla_k(z) = \\sum_j \\alpha_j(z) \\nabla_k n_j(z) + \\sum_j \\beta_j(z) \\nabla_k M(x_j(z))$$</p>\n<p>Therefore, the notebook idea blended solution can be formulated as a kernel function:</p>\n<p>$$K_j(z,z') = \\nabla_k(z) x_j(z')$$</p>\n<p>To optimize the kernel function, we can introduce a kernel matrix of all the kernels and their respective gradients for each notebook-idea-blended solution.</p>\n<p>$$\\hat{K}_{jk}(z,z') = \\nabla_k(z) x_j(z')$$</p>\n<h4>Past solution influence</h4>\n<p>Historical influence can sometime be attributed to some public submissions. Also with the recent social activity made popular by covid lockdown elimination this type of influence becomes more and more of a factor. Thus, from second law of medaldynamics we can assume that just like entropy, ideas flow from low medal density areas to high medal density areas, so we can define the past idea-solution as:</p>\n<p>$$n_j(z) = \\sum_i \\nabla_{i}(z)$$</p>\n<p>Where <em>𝛼</em> is the coefficient for idea divergence <em>(The second law of medaldynamics)</em> and <em>𝛽</em> is the past idea-solution.</p>\n<p>We can therefore define the <strong>emergence of creativity</strong> as the process of extracting the solution that is orthogonal to past influence.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1828038,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "06/21/2022 13:03:12",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> </p>\n<p>Er, wow! I guess that is why you are called \"The Devastator\"… 😃</p>\n<p>That said, somewhat more than your \"<em>Regularized Stochastic Leaderboard Descent</em>\"(!) I am sort of reminded of the <a href=\"https://en.wikipedia.org/wiki/Genetic_algorithm\" target=\"_blank\">genetic algorithms</a> that were very much the fashion some 20-30 years ago….</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1828393,
          "author_name": "karandora",
          "author_url": "",
          "post_date": "06/21/2022 16:45:49",
          "content": "<p>You should get likes simply for attempting to write this!!🙌</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1828408,
          "author_name": "thedevastator",
          "author_url": "",
          "post_date": "06/21/2022 17:02:09",
          "content": "<p>Thank you, Thank you 😁</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1832792,
          "author_name": "saurabhbagchi",
          "author_url": "",
          "post_date": "06/25/2022 11:54:52",
          "content": "<p>Great work <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a>, upvoted :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1828351,
      "author_name": "susnato",
      "author_url": "",
      "post_date": "06/21/2022 16:20:54",
      "content": "<p>Kagglers…Assemble.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1836195,
          "author_name": "mohammadrahmati",
          "author_url": "",
          "post_date": "06/28/2022 12:50:06",
          "content": "<p>nice one…!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1828405,
      "author_name": "christoffer",
      "author_url": "",
      "post_date": "06/21/2022 16:58:26",
      "content": "<p>IIRC, in the past, organizers have created aggregate models from submissions after the end of some competitions. I don’t remember which competititions but it might have been some of the basketball or american football ones.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1828419,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "06/21/2022 17:11:18",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/christoffer\" target=\"_blank\">@christoffer</a> </p>\n<p>That is interesting. Although from an organizers point of view it is worth distinguishing between ensembling a diverse set of models using the provided code (as per the <a href=\"https://web.archive.org/web/20160304031055/http://mlwave.com/kaggle-ensembling-guide/\" target=\"_blank\">\"Kaggle Ensembling Guide\"</a>) which can (could? <a href=\"https://www.wired.com/2012/04/netflix-prize-costs/\" target=\"_blank\">\"<em>Netflix Never Used Its $1 Million Algorithm Due To Engineering Costs</em>\"</a>) be beneficial to the organizers, as opposed to the blending of raw <code>submission.csv</code> files together, which may also do well on a kaggle  competition, but on the other hand provide little (basically nothing) of use to the organizers, as they are solutions specific to the dataset provided. (PS: I am a little surprised that this is not a \"Code\" competition).</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1829214,
          "author_name": "edwardakalarrywelch",
          "author_url": "",
          "post_date": "06/22/2022 12:55:47",
          "content": "<p>Thanks for that link.  I've been learning much more about ensembles and missed that resource.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1829985,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "06/23/2022 04:54:26",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/edwardakalarrywelch\" target=\"_blank\">@edwardakalarrywelch</a> </p>\n<p>Motivated by your comment I have just posted the link as a Topic. As this is the first big \"Featured\" non-forecasting tabular competition in quite a while, and could hit 4-5k participants, and there will be quite a few people who have not seen \"kaggle ensembling\" in the text books. Furthermore, the site MLwave went down about a year ago, but thankfully the content was rescued by the Wayback Machine. </p>\n<p>All the best,<br>\ncarl</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1829757,
      "author_name": "notabene",
      "author_url": "",
      "post_date": "06/22/2022 22:58:44",
      "content": "<p>Beautifully written. I scout the site for technical insights but, occasionally, I stumble on such meta posts which make the process not just useful, but also pleasurable.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1830437,
      "author_name": "heyspaceturtle",
      "author_url": "",
      "post_date": "06/23/2022 12:49:08",
      "content": "<p>People at AMEX were probably like: <br>\n— Hey we need an insane model for this business problem…<br>\n— Lets just fit a Kaggle Random Forest.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1830440,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "06/23/2022 12:53:17",
          "content": "<p>Hola <a href=\"https://www.kaggle.com/heyspaceturtle\" target=\"_blank\">@heyspaceturtle</a> </p>\n<p>Maybe, but it takes 3 months to run, and doesn't come cheap with $100,000 for prize money and an unknown amount in compute + storage…</p>\n<p>That said, credit default could cost somebody like American Express $M each year, so even a slightly better model could be worth the outlay.</p>\n<p>Un saludo muy codrial,<br>\ncarl</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1830464,
          "author_name": "heyspaceturtle",
          "author_url": "",
          "post_date": "06/23/2022 13:06:50",
          "content": "<p>True… but they still get the chance to find talented data scientists for their team :) </p>\n<p>Saludos! </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1827527": "Just a shower thought whilst running code; although ostensibly each of us are working alone to win this competition, we are in some way forming the \"Random Kaggle estimator\".\n\nEach team is a \"weak\" learner (not meant in any pejorative sense) akin to a decision tree estimator. However, collectively we are like the strong Random Forest leading to an almost optimal solution, using the Public Leaderboard as our objective function. We are very much guided by the work of others, either directly (for example forking and tweaking public notebooks) or indirectly (ideas as to what to explore further, or what others have already tried and do not seem to work, or even simply observing a sudden jump in the top LB score indicating that there is some *magic* to be found). The ideas and code do not even have to originate from this competition, but maybe are from a similar kaggle competition in the (perhaps distant) past, and some of the more recent Tabular Playground Series write-ups are astounding and truly inspiring!\n\nIt is this 'human-in-the-loop' meta-estimator that can be found only on kaggle! \n\nAnyway, I think my script has just finished, so back to work...\n\nGood luck to all; the better you do, the better everyone does!",
    "1827661": "if all team can made their submission csv file public .... then it will be interesting",
    "1827673": "Dear @hengck23 \n\nAmusing idea! One thing about stand-alone `submission.csv` files is that they are a bit like sausages; you don't know what went into making them. Indeed perhaps all closed-source blending notebooks should even come with a *caveat emptor*. For example, they must clearly say on the packet whether they have   [`B_29`](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328756) added into the mix or not 😃\n\nAll the best,\ncarl",
    "1827784": "We are all part of a big ensemble.",
    "1827816": "Hi @carlmcbrideellis I would be curious if there is some gain in averaging lots of score or if it is somehow capped. I would be also curious of innovative weighting scheme that could be used. I know that some market finance techniques (Hierarchical Risk Parity) was proposed.",
    "1827819": "Dear @lucasmorin \n\nI am completely unfamiliar with \"Hierarchical Risk Parity\", do you have any links?\n\nAll the best,\ncarl",
    "1827849": "well at least `B_29` is at the bottom of the list in terms how useful it is: )",
    "1827856": "Dear @raddar \n\nOnly for the training data and the Public chunk; ¿what happens when it 'kicks-in' on the Private data is the question?\n\nThat said, if the estimators they use employ just a smidgen of regularization (*i.e.* L2, or perhaps better still L1) they probably  create a model that diminishes (L2) or disregards (L1) the feature  `B_29` anyway...\n\nAll the best,\ncarl",
    "1827938": "HRP Paper: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2708678\nHRP Slides: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2713516\n\nI've seen the idea of applying portfolio selection approach to model ensembling as the problem are relatively similar (can't find a good ressource right now).",
    "1828000": "### Kaggle Random Walk\n##### Regularized Stochastic Leaderboard Descent\n_____\n\n### Abstract\nWe demonstrate the real work behind winning any Kaggle competition is actually accomplished through a human-computer collaboration. Teams are treated as individual weak learners who compete against one another by deriving optimized estimation solutions on the public leaderboard. Mutually and unconsciously, competitors in turn capitalize on the strengths and mitigate the weaknesses of other competitors by re-reporting their findings, when necessary, and ultimately developing their own solutions. By assembling the individual solutions together, we obtain an ideal solution, which is almost optimal. This \"Random Kaggle estimator\" enables us to demonstrate how humans, as part of an ML system, significantly contribute to its overall ability to solve complex problems. The more competitors, the more distinct differences there are between them (a min. two requisite), and the greater the leaderboard fluctuations, the faster and better the meta-estimation solution can be developed.\n\nFurthermore, we define a measure of the degree to which random actions by humans are required to accomplish a task relative to the task's complexity - \"Stochastic Leaderboard Descent\". We will demonstrate that the more complex a problem-solving task requires a greater number of independent teams working in unison to accomplish the ML task, with increasing importance placed on each team as the degree of stochasticity increases. We also show that although some of the ideas and code do not originate from the same competition (or are from the same competition at a different time), there are still ideas that we can not attribute to being influenced by any specific historical source.\n\nSuggesting that such ideas result from pure creativity and are driven by intuition, a human strength that extends far beyond ML tasks.\n\n#### Introduction\n\nIn the world of Kaggle, very often only the top spot on the leaderboard matters. On the other hand, once the leaderboard is released, it becomes a valuable tool for researchers to learn from and obtain new insights, especially in competitions with a sufficient number of competitors.\n\nIn this paper, we show that there is much more to leaderboards than the winner's name. Using a recent Kaggle competition as an example, we illustrate that competitors are able to adapt to one another and progressively improve their leaderboard score by benefiting from the cumulative effect of multiple contributions. We formulate a \"Random Kaggle estimator\" which, although simple, is able to reproduce the leaderboard evolution curve and we derive some insights from this model.\n\nWe then demonstrate how the leaderboard evolution curve itself can be used to define an optimal estimation model, which we call the \"stochastic leaderboard descent\".\n\n### Preliminaries\n\n#### Generalized Blending Theory\n\n**Definition 1: Solution Blending**\n> The process of assembling the predictions of multiple solutions in order to produce a new and improved estimate of the target value.\n\nTherefore, it follows from Definition 1 that a blended solution is the combination of multiple solutions that can be assembled to produce a more accurate estimate of the target value. As the blended solution is a combination of multiple solutions, it is a collaborative ensemble of solutions. Therefore, we can define a blended solution as follows:\n\n$$\\beta(x_i) = \\sum_j \\alpha_j(x_i) x_j(x_i)$$\n\nWhere *β* is the blended solution, *𝛼* is the solution weight and *x_j(x_i)* is the prediction for the training instance *x_i*.\n\n**Ideas Blending**\n\nBefore we define we must first define *Idea*:\n\nData scientists come up with a metric and then optimize that metric.\n\nAs we can see the information of the target outcome is linked to both the solution and the methods used to optimize the solution. We can represent a solution as follows:\n\n$$x_i(z) = \\Phi_i(z)\\phi_i(z)$$\n\nWhere *z* is the data, *Φ* is the encoder (Optimization method) and 𝜙 is the decoder (model), the feature-based solution.\n\nLastly, we are able to formulate the idea of data scientist as:\n\n$$y_i(z) = M(x_i(z))$$\n\nWhere *y_i(z)* is the data science outcome and *M* is the information linking the respective solution to the metric.\n\n**Definition 2: Ideas Blending**\n> The process of assembling the predictions of multiple ideas in order to produce a new and improved estimate of the target value.\n\nFrom Definition 1, we can define *Ideas Blending* as:\n\n$$y_i = \\arg \\min_{\\beta(z)} \\left[\\sum_j \\alpha_j y_j(z)\\right]$$\n\nSince the scientists try to blend ideas in order to minimize his leaderboard position.\n\n**Definition 3: Stochastic Leaderboard Convergence (SLC)**\n> A Stochastic Leaderboard Convergence is a process of obtaining an improved estimate of the target value by re-estimating the target value after a new public leaderboard score has been reported.\n\nIt follows from Definition 1 that a Stochastic Leaderboard Convergence (SLC) is a process of obtaining an improved estimate of the target value by re-estimating the target value after a new public leaderboard score has been reported. The idea behind SLC is that the more estimates of the target value are reported on the public leaderboard, the better the estimate of the target value will be.\n\nWe can therefore formulate the Stochastic Leaderboard Convergence (SLC) as follows:\n\n$$\\hat{y} = \\arg \\min_{\\beta(x_i)} \\left[\\sum_i \\left[\\sum_j \\alpha_j(x_i) x_j(x_i) - y_i\\right]^2\\right]$$\n\n**Definition 4: Public Leaderboard Density**\n> The density of the number of public leaderboard scores reported.\n\nThe Public Leaderboard Density (PLD) is the number of reported public leaderboard scores per unit area (competition timeframe). We can formulate the PLD as follows:\n\n$$PLD = \\frac{N}{T}$$\n\nWhere *N* is the number of public leaderboard scores reported, and *T* is the competition timeframe.\n\n**Definition 4: Stochastic Leaderboard Descent (SLD)**\n> A Stochastic Leaderboard Descent is a process of obtaining an improved estimate of the target value by taking the current public leaderboard score as the starting point and applying stochastic gradient descent to obtain a better estimate of the target value.\n\n#### Leaderboard Monte Carlo Estimation\n\n**Bayesian Blending**\n\nBayesian Blending is a process of obtaining an improved estimate of the target value by taking the current top public solution as the starting point and applying stochastic notebook descent to obtain a better estimate of the target value.\nSince there are: \n\n- **Solutions** released as notebooks\n- **Ideas** released on the discussions forums\n\nWe propse to define the blended solution as a linear combination of both a `solution term` and an `idea term`:\n\n$$\\omega(x_i) = \\sum_j \\alpha_j n_j(z) + \\sum_k \\beta_k M(x_k(z))$$\n\nWhere *n_j* is the notebook's code used (function of the data *z*) and *M(x_k(z))* again is the information extracted from an idea *x_i* (function of the competition's data). The weights *𝛼* and *𝛽* are the probabilistic weight of the notebooks and ideas.\n\nWe therefore can define the direction of the kaggle bayesian blending process as the gradient of the notebook-idea-blended solution.\n\n$$\\nabla_k(z) = \\sum_j \\alpha_j(z) \\nabla_k n_j(z) + \\sum_j \\beta_j(z) \\nabla_k M(x_j(z))$$\n\nTherefore, the notebook idea blended solution can be formulated as a kernel function:\n\n$$K_j(z,z') = \\nabla_k(z) x_j(z')$$\n\nTo optimize the kernel function, we can introduce a kernel matrix of all the kernels and their respective gradients for each notebook-idea-blended solution.\n\n$$\\hat{K}_{jk}(z,z') = \\nabla_k(z) x_j(z')$$\n\n#### Past solution influence\n\nHistorical influence can sometime be attributed to some public submissions. Also with the recent social activity made popular by covid lockdown elimination this type of influence becomes more and more of a factor. Thus, from second law of medaldynamics we can assume that just like entropy, ideas flow from low medal density areas to high medal density areas, so we can define the past idea-solution as:\n\n$$n_j(z) = \\sum_i \\nabla_{i}(z)$$\n\nWhere *𝛼* is the coefficient for idea divergence *(The second law of medaldynamics)* and *𝛽* is the past idea-solution.\n\nWe can therefore define the **emergence of creativity** as the process of extracting the solution that is orthogonal to past influence.",
    "1828038": "Dear @thedevastator \n\nEr, wow! I guess that is why you are called \"The Devastator\"... 😃\n\nThat said, somewhat more than your \"*Regularized Stochastic Leaderboard Descent*\"(!) I am sort of reminded of the [genetic algorithms](https://en.wikipedia.org/wiki/Genetic_algorithm) that were very much the fashion some 20-30 years ago....\n\nAll the best,\ncarl",
    "1828351": "Kagglers...Assemble.",
    "1828393": "You should get likes simply for attempting to write this!!🙌",
    "1828405": "IIRC, in the past, organizers have created aggregate models from submissions after the end of some competitions. I don’t remember which competititions but it might have been some of the basketball or american football ones.",
    "1828408": "Thank you, Thank you 😁",
    "1828419": "Dear @christoffer \n\nThat is interesting. Although from an organizers point of view it is worth distinguishing between ensembling a diverse set of models using the provided code (as per the [\"Kaggle Ensembling Guide\"](https://web.archive.org/web/20160304031055/http://mlwave.com/kaggle-ensembling-guide/)) which can (could? [\"*Netflix Never Used Its $1 Million Algorithm Due To Engineering Costs*\"](https://www.wired.com/2012/04/netflix-prize-costs/)) be beneficial to the organizers, as opposed to the blending of raw `submission.csv` files together, which may also do well on a kaggle  competition, but on the other hand provide little (basically nothing) of use to the organizers, as they are solutions specific to the dataset provided. (PS: I am a little surprised that this is not a \"Code\" competition).\n\nAll the best,\ncarl",
    "1829214": "Thanks for that link.  I've been learning much more about ensembles and missed that resource.",
    "1829757": "Beautifully written. I scout the site for technical insights but, occasionally, I stumble on such meta posts which make the process not just useful, but also pleasurable.",
    "1829985": "Dear @edwardakalarrywelch \n\nMotivated by your comment I have just posted the link as a Topic. As this is the first big \"Featured\" non-forecasting tabular competition in quite a while, and could hit 4-5k participants, and there will be quite a few people who have not seen \"kaggle ensembling\" in the text books. Furthermore, the site MLwave went down about a year ago, but thankfully the content was rescued by the Wayback Machine. \n\nAll the best,\ncarl",
    "1830437": "People at AMEX were probably like: \n— Hey we need an insane model for this business problem...\n— Lets just fit a Kaggle Random Forest.",
    "1830440": "Hola @heyspaceturtle \n\nMaybe, but it takes 3 months to run, and doesn't come cheap with $100,000 for prize money and an unknown amount in compute + storage...\n\nThat said, credit default could cost somebody like American Express $M each year, so even a slightly better model could be worth the outlay.\n\nUn saludo muy codrial,\ncarl",
    "1830464": "True... but they still get the chance to find talented data scientists for their team :) \n\nSaludos!",
    "1832792": "Great work @thedevastator, upvoted :)",
    "1836195": "nice one...!"
  },
  "source": "meta"
}