{
  "id": 74540,
  "title": "Trouble with ensemble",
  "url": "/competitions/PLAsTiCC-2018/discussion/74540",
  "author_name": "",
  "post_date": "2018-12-13T09:05:39.513948800Z",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I used the 'blend them' public kernel as a template and got a small boost with an ensemble of my various tree-based models.  However I knew I needed a really different approach so I worked feverishly to get a decent (1.056) neural-net model to blend with my lgbm model (1.010).</p>\n\n<p>I waited with eager anticipation only to see 1.445 come back.  I've been through the dataFrames forward and backward (I assumed maybe the sort order mattered or something else along those lines).  I'm really perplexed.  Has anyone else encountered unexpected pitfalls while trying to ensemble?</p>",
  "messages": [
    {
      "id": "438204",
      "postDate": "12/13/2018 09:05:39",
      "content": "<p>I used the 'blend them' public kernel as a template and got a small boost with an ensemble of my various tree-based models.  However I knew I needed a really different approach so I worked feverishly to get a decent (1.056) neural-net model to blend with my lgbm model (1.010).</p>\n\n<p>I waited with eager anticipation only to see 1.445 come back.  I've been through the dataFrames forward and backward (I assumed maybe the sort order mattered or something else along those lines).  I'm really perplexed.  Has anyone else encountered unexpected pitfalls while trying to ensemble?</p>",
      "rawMarkdown": "I used the 'blend them' public kernel as a template and got a small boost with an ensemble of my various tree-based models.  However I knew I needed a really different approach so I worked feverishly to get a decent (1.056) neural-net model to blend with my lgbm model (1.010).\n\nI waited with eager anticipation only to see 1.445 come back.  I've been through the dataFrames forward and backward (I assumed maybe the sort order mattered or something else along those lines).  I'm really perplexed.  Has anyone else encountered unexpected pitfalls while trying to ensemble?",
      "votes": null
    },
    {
      "id": "438214",
      "postDate": "12/13/2018 09:28:06",
      "content": "<p>I wouldn't have expected a blend of two decent models to perform so poorly. Good luck finding the problem.</p>",
      "rawMarkdown": "I wouldn't have expected a blend of two decent models to perform so poorly. Good luck finding the problem.",
      "votes": null
    },
    {
      "id": "438239",
      "postDate": "12/13/2018 10:20:30",
      "content": "<p>usually I use lightGBM, knn and deep learning model for ensembling ( for regression problems) as it captures good variance. Multi-class classification is trickier ! My single model is 1.029.. didnt try ensembling yet.. Any suggestions?  </p>",
      "rawMarkdown": "usually I use lightGBM, knn and deep learning model for ensembling ( for regression problems) as it captures good variance. Multi-class classification is trickier ! My single model is 1.029.. didnt try ensembling yet.. Any suggestions?",
      "votes": null
    },
    {
      "id": "438273",
      "postDate": "12/13/2018 11:53:57",
      "content": "<p>Are you sure you are averaging with the submissions in the same order - I ask because you mention sorting</p>",
      "rawMarkdown": "Are you sure you are averaging with the submissions in the same order - I ask because you mention sorting",
      "votes": null
    },
    {
      "id": "438278",
      "postDate": "12/13/2018 11:58:23",
      "content": "<ol>\n<li><p>Do you see improvement using the same way you combine test predictions, but using out of fold predictions on train data?</p>\n\n<ol><li>As Scirpus said, are you sure you don't have indexing issues?</li></ol></li>\n</ol>",
      "rawMarkdown": "1. Do you see improvement using the same way you combine test predictions, but using out of fold predictions on train data?\n\n2. As Scirpus said, are you sure you don't have indexing issues?",
      "votes": null
    },
    {
      "id": "438615",
      "postDate": "12/14/2018 00:10:35",
      "content": "<p>I think I had indexing issues:</p>\n\n<p>import random\ndef validateBlend(bdf, dfs, wts, npc=1000):</p>\n\n<pre><code>recs=bdf.shape[0]\ncols=len(bdf.columns)\nfailures=0\nfor i in range(npc):\n    record=random.randint(0,recs-1)\n    feat=bdf.columns[random.randint(0, cols-1)]\n\n\n    shouldBe=0\n    isVal=bdf.loc[record, feat]\n    for j in range(len(wts)):\n\n        shouldBe += dfs[j].loc[record, feat] * wts[j]\n    shouldBe=round(shouldBe,6)\n    isVal=round(isVal,6)\n    if shouldBe==isVal:\n        #print('match')\n        dosomething=False\n    else:\n        print(record)\n        print(feat)\n        print('should be: ' + str(shouldBe))\n        print('is: ' + str(isVal))\n        failures+=1\n\nprint('total failures: ' + str(failures) + ' out of ' + str(npc) +' sample points')\nreturn failures\n</code></pre>\n\n<p>failures=validateBlend(dfBlend, dfs, wts)    </p>",
      "rawMarkdown": "I think I had indexing issues:\n\n\nimport random\ndef validateBlend(bdf, dfs, wts, npc=1000):\n    \n    recs=bdf.shape[0]\n    cols=len(bdf.columns)\n    failures=0\n    for i in range(npc):\n        record=random.randint(0,recs-1)\n        feat=bdf.columns[random.randint(0, cols-1)]\n        \n        \n        shouldBe=0\n        isVal=bdf.loc[record, feat]\n        for j in range(len(wts)):\n            \n            shouldBe += dfs[j].loc[record, feat] * wts[j]\n        shouldBe=round(shouldBe,6)\n        isVal=round(isVal,6)\n        if shouldBe==isVal:\n            #print('match')\n            dosomething=False\n        else:\n            print(record)\n            print(feat)\n            print('should be: ' + str(shouldBe))\n            print('is: ' + str(isVal))\n            failures+=1\n            \n    print('total failures: ' + str(failures) + ' out of ' + str(npc) +' sample points')\n    return failures\n            \n    \nfailures=validateBlend(dfBlend, dfs, wts)",
      "votes": null
    },
    {
      "id": "440050",
      "postDate": "12/17/2018 00:44:29",
      "content": "<p><a href=\"/jimpsull\">@jimpsull</a> could you link to the public kernel you were talking about ('blend them')?</p>",
      "rawMarkdown": "jimpsull could you link to the public kernel you were talking about ('blend them')?",
      "votes": null
    },
    {
      "id": "440073",
      "postDate": "12/17/2018 02:00:56",
      "content": "<p>We've had trouble with lots of things here ;)</p>",
      "rawMarkdown": "We've had trouble with lots of things here ;)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 438214,
      "author_name": "philipwahrlich",
      "author_url": "",
      "post_date": "12/13/2018 09:28:06",
      "content": "<p>I wouldn't have expected a blend of two decent models to perform so poorly. Good luck finding the problem.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 438239,
      "author_name": "prashanththangavel",
      "author_url": "",
      "post_date": "12/13/2018 10:20:30",
      "content": "<p>usually I use lightGBM, knn and deep learning model for ensembling ( for regression problems) as it captures good variance. Multi-class classification is trickier ! My single model is 1.029.. didnt try ensembling yet.. Any suggestions?  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 438273,
      "author_name": "scirpus",
      "author_url": "",
      "post_date": "12/13/2018 11:53:57",
      "content": "<p>Are you sure you are averaging with the submissions in the same order - I ask because you mention sorting</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 438278,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "12/13/2018 11:58:23",
      "content": "<ol>\n<li><p>Do you see improvement using the same way you combine test predictions, but using out of fold predictions on train data?</p>\n\n<ol><li>As Scirpus said, are you sure you don't have indexing issues?</li></ol></li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 438615,
      "author_name": "jimpsull",
      "author_url": "",
      "post_date": "12/14/2018 00:10:35",
      "content": "<p>I think I had indexing issues:</p>\n\n<p>import random\ndef validateBlend(bdf, dfs, wts, npc=1000):</p>\n\n<pre><code>recs=bdf.shape[0]\ncols=len(bdf.columns)\nfailures=0\nfor i in range(npc):\n    record=random.randint(0,recs-1)\n    feat=bdf.columns[random.randint(0, cols-1)]\n\n\n    shouldBe=0\n    isVal=bdf.loc[record, feat]\n    for j in range(len(wts)):\n\n        shouldBe += dfs[j].loc[record, feat] * wts[j]\n    shouldBe=round(shouldBe,6)\n    isVal=round(isVal,6)\n    if shouldBe==isVal:\n        #print('match')\n        dosomething=False\n    else:\n        print(record)\n        print(feat)\n        print('should be: ' + str(shouldBe))\n        print('is: ' + str(isVal))\n        failures+=1\n\nprint('total failures: ' + str(failures) + ' out of ' + str(npc) +' sample points')\nreturn failures\n</code></pre>\n\n<p>failures=validateBlend(dfBlend, dfs, wts)    </p>",
      "votes": null,
      "replies": [
        {
          "id": 440050,
          "author_name": "authman",
          "author_url": "",
          "post_date": "12/17/2018 00:44:29",
          "content": "<p><a href=\"/jimpsull\">@jimpsull</a> could you link to the public kernel you were talking about ('blend them')?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 440073,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "12/17/2018 02:00:56",
      "content": "<p>We've had trouble with lots of things here ;)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "438204": "I used the 'blend them' public kernel as a template and got a small boost with an ensemble of my various tree-based models.  However I knew I needed a really different approach so I worked feverishly to get a decent (1.056) neural-net model to blend with my lgbm model (1.010).\n\nI waited with eager anticipation only to see 1.445 come back.  I've been through the dataFrames forward and backward (I assumed maybe the sort order mattered or something else along those lines).  I'm really perplexed.  Has anyone else encountered unexpected pitfalls while trying to ensemble?",
    "438214": "I wouldn't have expected a blend of two decent models to perform so poorly. Good luck finding the problem.",
    "438239": "usually I use lightGBM, knn and deep learning model for ensembling ( for regression problems) as it captures good variance. Multi-class classification is trickier ! My single model is 1.029.. didnt try ensembling yet.. Any suggestions?",
    "438273": "Are you sure you are averaging with the submissions in the same order - I ask because you mention sorting",
    "438278": "1. Do you see improvement using the same way you combine test predictions, but using out of fold predictions on train data?\n\n2. As Scirpus said, are you sure you don't have indexing issues?",
    "438615": "I think I had indexing issues:\n\n\nimport random\ndef validateBlend(bdf, dfs, wts, npc=1000):\n    \n    recs=bdf.shape[0]\n    cols=len(bdf.columns)\n    failures=0\n    for i in range(npc):\n        record=random.randint(0,recs-1)\n        feat=bdf.columns[random.randint(0, cols-1)]\n        \n        \n        shouldBe=0\n        isVal=bdf.loc[record, feat]\n        for j in range(len(wts)):\n            \n            shouldBe += dfs[j].loc[record, feat] * wts[j]\n        shouldBe=round(shouldBe,6)\n        isVal=round(isVal,6)\n        if shouldBe==isVal:\n            #print('match')\n            dosomething=False\n        else:\n            print(record)\n            print(feat)\n            print('should be: ' + str(shouldBe))\n            print('is: ' + str(isVal))\n            failures+=1\n            \n    print('total failures: ' + str(failures) + ' out of ' + str(npc) +' sample points')\n    return failures\n            \n    \nfailures=validateBlend(dfBlend, dfs, wts)",
    "440050": "jimpsull could you link to the public kernel you were talking about ('blend them')?",
    "440073": "We've had trouble with lots of things here ;)"
  },
  "source": "meta"
}