{
  "id": 197145,
  "title": "Sqlite3 in Prediction Loop",
  "url": "/competitions/riiid-test-answer-prediction/discussion/197145",
  "author_name": "",
  "post_date": "2020-11-14T16:26:19.192033800Z",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F781502%2F16d1336aab2164f4c7576a5ad64bd17a%2F2020-11-14_8-29-09.png?generation=1605371401972691&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://www.kaggle.com/calebeverett/riiid-submit\" target=\"_blank\">This</a> was looking super <a href=\"https://blog.thedataincubator.com/2018/05/sqlite-vs-pandas-performance-benchmarks/\" target=\"_blank\">promising</a>, making its way through 2.5 million records in a little over three hours, including a select from an 80 million record user-content table, a 300k million record user table and a 13k records questions table.</p>\n<p>In addition to being fast, I really liked the idea of being able to add features and incorporate them into the prediction loop with a simple update of a few sql statements.</p>\n<p>Unfortunately, it is not making its way through the hidden test set. I've been banging my head on it for the last several days and thought I'd see if anybody else has any ideas that might salvage it before moving on.</p>\n<p>I've isolated it to the select statement. My current hypothesis is that the median batch size in the test set is really small. This approach crushes it on large batch sizes - close to 2k per second at batch sizes of 1000 - but suffers when they are small. If that is the case, then this approach won't work, but perhaps something else is causing it to bonk.</p>",
  "messages": [
    {
      "id": "1078326",
      "postDate": "11/14/2020 16:26:19",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F781502%2F16d1336aab2164f4c7576a5ad64bd17a%2F2020-11-14_8-29-09.png?generation=1605371401972691&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://www.kaggle.com/calebeverett/riiid-submit\" target=\"_blank\">This</a> was looking super <a href=\"https://blog.thedataincubator.com/2018/05/sqlite-vs-pandas-performance-benchmarks/\" target=\"_blank\">promising</a>, making its way through 2.5 million records in a little over three hours, including a select from an 80 million record user-content table, a 300k million record user table and a 13k records questions table.</p>\n<p>In addition to being fast, I really liked the idea of being able to add features and incorporate them into the prediction loop with a simple update of a few sql statements.</p>\n<p>Unfortunately, it is not making its way through the hidden test set. I've been banging my head on it for the last several days and thought I'd see if anybody else has any ideas that might salvage it before moving on.</p>\n<p>I've isolated it to the select statement. My current hypothesis is that the median batch size in the test set is really small. This approach crushes it on large batch sizes - close to 2k per second at batch sizes of 1000 - but suffers when they are small. If that is the case, then this approach won't work, but perhaps something else is causing it to bonk.</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F781502%2F16d1336aab2164f4c7576a5ad64bd17a%2F2020-11-14_8-29-09.png?generation=1605371401972691&alt=media)\n\n[This](https://www.kaggle.com/calebeverett/riiid-submit) was looking super [promising](https://blog.thedataincubator.com/2018/05/sqlite-vs-pandas-performance-benchmarks/), making its way through 2.5 million records in a little over three hours, including a select from an 80 million record user-content table, a 300k million record user table and a 13k records questions table.\n\nIn addition to being fast, I really liked the idea of being able to add features and incorporate them into the prediction loop with a simple update of a few sql statements.\n\nUnfortunately, it is not making its way through the hidden test set. I've been banging my head on it for the last several days and thought I'd see if anybody else has any ideas that might salvage it before moving on.\n\nI've isolated it to the select statement. My current hypothesis is that the median batch size in the test set is really small. This approach crushes it on large batch sizes - close to 2k per second at batch sizes of 1000 - but suffers when they are small. If that is the case, then this approach won't work, but perhaps something else is causing it to bonk.",
      "votes": null
    },
    {
      "id": "1078593",
      "postDate": "11/15/2020 04:14:19",
      "content": "<p>Others have noted elsewhere (<a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124#1056282\" target=\"_blank\">here for example</a>) that the batch sizes are likely very small, median of like 10-20 or so</p>",
      "rawMarkdown": "Others have noted elsewhere ([here for example](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124#1056282)) that the batch sizes are likely very small, median of like 10-20 or so",
      "votes": null
    },
    {
      "id": "1079200",
      "postDate": "11/15/2020 18:36:32",
      "content": "<p>Thanks, I got it sorted.</p>",
      "rawMarkdown": "Thanks, I got it sorted.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1078593,
      "author_name": "brandenkmurray",
      "author_url": "",
      "post_date": "11/15/2020 04:14:19",
      "content": "<p>Others have noted elsewhere (<a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124#1056282\" target=\"_blank\">here for example</a>) that the batch sizes are likely very small, median of like 10-20 or so</p>",
      "votes": null,
      "replies": [
        {
          "id": 1079200,
          "author_name": "calebeverett",
          "author_url": "",
          "post_date": "11/15/2020 18:36:32",
          "content": "<p>Thanks, I got it sorted.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1078326": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F781502%2F16d1336aab2164f4c7576a5ad64bd17a%2F2020-11-14_8-29-09.png?generation=1605371401972691&alt=media)\n\n[This](https://www.kaggle.com/calebeverett/riiid-submit) was looking super [promising](https://blog.thedataincubator.com/2018/05/sqlite-vs-pandas-performance-benchmarks/), making its way through 2.5 million records in a little over three hours, including a select from an 80 million record user-content table, a 300k million record user table and a 13k records questions table.\n\nIn addition to being fast, I really liked the idea of being able to add features and incorporate them into the prediction loop with a simple update of a few sql statements.\n\nUnfortunately, it is not making its way through the hidden test set. I've been banging my head on it for the last several days and thought I'd see if anybody else has any ideas that might salvage it before moving on.\n\nI've isolated it to the select statement. My current hypothesis is that the median batch size in the test set is really small. This approach crushes it on large batch sizes - close to 2k per second at batch sizes of 1000 - but suffers when they are small. If that is the case, then this approach won't work, but perhaps something else is causing it to bonk.",
    "1078593": "Others have noted elsewhere ([here for example](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124#1056282)) that the batch sizes are likely very small, median of like 10-20 or so",
    "1079200": "Thanks, I got it sorted."
  },
  "source": "meta"
}