{
  "id": 199719,
  "title": "classification approach",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/writeups/outrunner-classification-approach",
  "author_name": "",
  "post_date": "2020-11-27T02:35:47.312419800Z",
  "votes": 11,
  "comment_count": 5,
  "views": 0,
  "content": "<p>In the last week, I tried classification approach and it performed well or even better than my regression model.</p>\n<p>Due to time limitations, I just use K-Means on validation trajectories to get 1K target classes. In test, merge Top70 predicts to 3. Last day, I trained a 5K classes model and use Top400 to merge.</p>\n<p>Blend regression(three models) and classification(500, 1K, 5K classes) models to get public LB score improvement from 15 to 12.8</p>\n<p>My merge code:</p>\n<p>`from itertools import combinations</p>\n<p>def optimize_predicts(preds, topK=3, yp=1):<br>\n    pred3 = []<br>\n    comb = np.array((list(combinations(range(preds.shape[-2]), topK))))</p>\n<pre><code>for k in range(len(preds)):\n\n    p = preds[k,:,:-1]\n    y = preds[k,:,-1]\n\n    cost = pairwise_distances(p)*(y**yp)\n\n    labels = cost[comb[cost[comb].min(1).sum(-1).argmin()]].argmin(0)\n\n    pr=np.zeros((topK,101))\n\n    for i in range(topK):\n        sel = labels==i\n        pr[i,:100] = ((p[sel]*y[sel,None]).sum(0)/y[sel].sum())\n        pr[i,-1] = y[sel].sum()\n\n    pred3.append(pr)\n\nreturn np.array(pred3)`\n</code></pre>",
  "messages": [
    {
      "id": "1092592",
      "postDate": "11/27/2020 02:35:47",
      "content": "<p>In the last week, I tried classification approach and it performed well or even better than my regression model.</p>\n<p>Due to time limitations, I just use K-Means on validation trajectories to get 1K target classes. In test, merge Top70 predicts to 3. Last day, I trained a 5K classes model and use Top400 to merge.</p>\n<p>Blend regression(three models) and classification(500, 1K, 5K classes) models to get public LB score improvement from 15 to 12.8</p>\n<p>My merge code:</p>\n<p>`from itertools import combinations</p>\n<p>def optimize_predicts(preds, topK=3, yp=1):<br>\n    pred3 = []<br>\n    comb = np.array((list(combinations(range(preds.shape[-2]), topK))))</p>\n<pre><code>for k in range(len(preds)):\n\n    p = preds[k,:,:-1]\n    y = preds[k,:,-1]\n\n    cost = pairwise_distances(p)*(y**yp)\n\n    labels = cost[comb[cost[comb].min(1).sum(-1).argmin()]].argmin(0)\n\n    pr=np.zeros((topK,101))\n\n    for i in range(topK):\n        sel = labels==i\n        pr[i,:100] = ((p[sel]*y[sel,None]).sum(0)/y[sel].sum())\n        pr[i,-1] = y[sel].sum()\n\n    pred3.append(pr)\n\nreturn np.array(pred3)`\n</code></pre>",
      "rawMarkdown": "In the last week, I tried classification approach and it performed well or even better than my regression model.\n\nDue to time limitations, I just use K-Means on validation trajectories to get 1K target classes. In test, merge Top70 predicts to 3. Last day, I trained a 5K classes model and use Top400 to merge.\n\nBlend regression(three models) and classification(500, 1K, 5K classes) models to get public LB score improvement from 15 to 12.8\n\nMy merge code:\n\n`from itertools import combinations\n\ndef optimize_predicts(preds, topK=3, yp=1):\n    pred3 = []\n    comb = np.array((list(combinations(range(preds.shape[-2]), topK))))\n\n    for k in range(len(preds)):\n        \n        p = preds[k,:,:-1]\n        y = preds[k,:,-1]\n        \n        cost = pairwise_distances(p)*(y**yp)\n        \n        labels = cost[comb[cost[comb].min(1).sum(-1).argmin()]].argmin(0)\n        \n        pr=np.zeros((topK,101))\n        \n        for i in range(topK):\n            sel = labels==i\n            pr[i,:100] = ((p[sel]*y[sel,None]).sum(0)/y[sel].sum())\n            pr[i,-1] = y[sel].sum()\n            \n        pred3.append(pr)\n    \n    return np.array(pred3)`",
      "votes": null
    },
    {
      "id": "1092717",
      "postDate": "11/27/2020 05:37:37",
      "content": "<p>So just to make sure I understand, what you did is load in the future positions and then apply k-means to them and extract out the cluster means and then used a model to predict which one was closest to the correct trajectory in the training set. Then you once again grabbed the top-n trajectories determined by your classifier and then applied k-means again to reduce down to 3 trajectories?</p>",
      "rawMarkdown": "So just to make sure I understand, what you did is load in the future positions and then apply k-means to them and extract out the cluster means and then used a model to predict which one was closest to the correct trajectory in the training set. Then you once again grabbed the top-n trajectories determined by your classifier and then applied k-means again to reduce down to 3 trajectories?",
      "votes": null
    },
    {
      "id": "1092728",
      "postDate": "11/27/2020 05:48:53",
      "content": "<p>Yes 😊<br>\n(Your message must have at least 10 characters.)</p>",
      "rawMarkdown": "Yes 😊\n(Your message must have at least 10 characters.)",
      "votes": null
    },
    {
      "id": "1092817",
      "postDate": "11/27/2020 07:34:30",
      "content": "<p>abc                  </p>",
      "rawMarkdown": "abc",
      "votes": null
    },
    {
      "id": "1092819",
      "postDate": "11/27/2020 07:35:10",
      "content": "<p><a href=\"https://www.kaggle.com/outrunner\" target=\"_blank\">@outrunner</a> <br>\nless than 10 characters is possible (see above)</p>",
      "rawMarkdown": "outrunner \nless than 10 characters is possible (see above)",
      "votes": null
    },
    {
      "id": "1092867",
      "postDate": "11/27/2020 08:26:09",
      "content": "<p>thanks <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 😊</p>",
      "rawMarkdown": "thanks @hengck23 😊",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1092717,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "11/27/2020 05:37:37",
      "content": "<p>So just to make sure I understand, what you did is load in the future positions and then apply k-means to them and extract out the cluster means and then used a model to predict which one was closest to the correct trajectory in the training set. Then you once again grabbed the top-n trajectories determined by your classifier and then applied k-means again to reduce down to 3 trajectories?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1092728,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "11/27/2020 05:48:53",
          "content": "<p>Yes 😊<br>\n(Your message must have at least 10 characters.)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1092817,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/27/2020 07:34:30",
          "content": "<p>abc                  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1092819,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/27/2020 07:35:10",
          "content": "<p><a href=\"https://www.kaggle.com/outrunner\" target=\"_blank\">@outrunner</a> <br>\nless than 10 characters is possible (see above)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1092867,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "11/27/2020 08:26:09",
          "content": "<p>thanks <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 😊</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1092592": "In the last week, I tried classification approach and it performed well or even better than my regression model.\n\nDue to time limitations, I just use K-Means on validation trajectories to get 1K target classes. In test, merge Top70 predicts to 3. Last day, I trained a 5K classes model and use Top400 to merge.\n\nBlend regression(three models) and classification(500, 1K, 5K classes) models to get public LB score improvement from 15 to 12.8\n\nMy merge code:\n\n`from itertools import combinations\n\ndef optimize_predicts(preds, topK=3, yp=1):\n    pred3 = []\n    comb = np.array((list(combinations(range(preds.shape[-2]), topK))))\n\n    for k in range(len(preds)):\n        \n        p = preds[k,:,:-1]\n        y = preds[k,:,-1]\n        \n        cost = pairwise_distances(p)*(y**yp)\n        \n        labels = cost[comb[cost[comb].min(1).sum(-1).argmin()]].argmin(0)\n        \n        pr=np.zeros((topK,101))\n        \n        for i in range(topK):\n            sel = labels==i\n            pr[i,:100] = ((p[sel]*y[sel,None]).sum(0)/y[sel].sum())\n            pr[i,-1] = y[sel].sum()\n            \n        pred3.append(pr)\n    \n    return np.array(pred3)`",
    "1092717": "So just to make sure I understand, what you did is load in the future positions and then apply k-means to them and extract out the cluster means and then used a model to predict which one was closest to the correct trajectory in the training set. Then you once again grabbed the top-n trajectories determined by your classifier and then applied k-means again to reduce down to 3 trajectories?",
    "1092728": "Yes 😊\n(Your message must have at least 10 characters.)",
    "1092817": "abc",
    "1092819": "outrunner \nless than 10 characters is possible (see above)",
    "1092867": "thanks @hengck23 😊"
  },
  "source": "meta"
}