{
  "id": 190439,
  "title": "Fake API generation from train set",
  "url": "/competitions/riiid-test-answer-prediction/discussion/190439",
  "author_name": "",
  "post_date": "2020-10-11T20:14:20.243818300Z",
  "votes": 13,
  "comment_count": 3,
  "views": 0,
  "content": "<p>As this competition pose a great dataflow optimization challenge, i wrote a kernel to generate a fake test set from the train set in a similar fashion as the real api from the competition, for debugging and optimisation purpose.</p>\n<p>A few differences to observe compared to the real api:</p>\n<ul>\n<li>the iterator need an index to be called</li>\n<li>you don't haveto call the function predict to go to the next batch</li>\n<li>the call is a bit faster than the real api, so you might want to add a delay to simulate a real call (my api : 1ms / call against about 10ms per call for the real one)</li>\n</ul>\n<p>The kernel can be found here : <a href=\"https://www.kaggle.com/rously/fake-api-generation-notebook-riid\" target=\"_blank\">https://www.kaggle.com/rously/fake-api-generation-notebook-riid</a><br>\nA preprocessed test set of about 2.4M lines can be found here:  <a href=\"https://www.kaggle.com/rously/fake-api-generation-riid-preprocessed\" target=\"_blank\">https://www.kaggle.com/rously/fake-api-generation-riid-preprocessed</a></p>\n<p>If you have any improvemement idea for this kernel, please feel free to leave in comments</p>",
  "messages": [
    {
      "id": "1046593",
      "postDate": "10/11/2020 20:14:20",
      "content": "<p>As this competition pose a great dataflow optimization challenge, i wrote a kernel to generate a fake test set from the train set in a similar fashion as the real api from the competition, for debugging and optimisation purpose.</p>\n<p>A few differences to observe compared to the real api:</p>\n<ul>\n<li>the iterator need an index to be called</li>\n<li>you don't haveto call the function predict to go to the next batch</li>\n<li>the call is a bit faster than the real api, so you might want to add a delay to simulate a real call (my api : 1ms / call against about 10ms per call for the real one)</li>\n</ul>\n<p>The kernel can be found here : <a href=\"https://www.kaggle.com/rously/fake-api-generation-notebook-riid\" target=\"_blank\">https://www.kaggle.com/rously/fake-api-generation-notebook-riid</a><br>\nA preprocessed test set of about 2.4M lines can be found here:  <a href=\"https://www.kaggle.com/rously/fake-api-generation-riid-preprocessed\" target=\"_blank\">https://www.kaggle.com/rously/fake-api-generation-riid-preprocessed</a></p>\n<p>If you have any improvemement idea for this kernel, please feel free to leave in comments</p>",
      "rawMarkdown": "As this competition pose a great dataflow optimization challenge, i wrote a kernel to generate a fake test set from the train set in a similar fashion as the real api from the competition, for debugging and optimisation purpose.\n\nA few differences to observe compared to the real api:\n- the iterator need an index to be called\n- you don't haveto call the function predict to go to the next batch\n- the call is a bit faster than the real api, so you might want to add a delay to simulate a real call (my api : 1ms / call against about 10ms per call for the real one)\n\n\nThe kernel can be found here : https://www.kaggle.com/rously/fake-api-generation-notebook-riid\nA preprocessed test set of about 2.4M lines can be found here:  https://www.kaggle.com/rously/fake-api-generation-riid-preprocessed\n\nIf you have any improvemement idea for this kernel, please feel free to leave in comments",
      "votes": null
    },
    {
      "id": "1046607",
      "postDate": "10/11/2020 20:36:31",
      "content": "<p>Thank you, I was just planning on doing something similar. How you measured \"100ms per call for the real one\"… or is this an approximate calculation?</p>",
      "rawMarkdown": "Thank you, I was just planning on doing something similar. How you measured \"100ms per call for the real one\"... or is this an approximate calculation?",
      "votes": null
    },
    {
      "id": "1046609",
      "postDate": "10/11/2020 20:40:28",
      "content": "<p><a href=\"https://www.kaggle.com/sapr3s\" target=\"_blank\">@sapr3s</a>  i made a mistake in my original post, it's more like 10ms per call instead of 100, to measure it i just made a few measurments by putting a %%time decorator while executing the API call :)</p>",
      "rawMarkdown": "sapr3s  i made a mistake in my original post, it's more like 10ms per call instead of 100, to measure it i just made a few measurments by putting a %%time decorator while executing the API call :)",
      "votes": null
    },
    {
      "id": "1047925",
      "postDate": "10/13/2020 04:14:12",
      "content": "<p>This is just perfect!</p>",
      "rawMarkdown": "This is just perfect!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1046607,
      "author_name": "sapr3s",
      "author_url": "",
      "post_date": "10/11/2020 20:36:31",
      "content": "<p>Thank you, I was just planning on doing something similar. How you measured \"100ms per call for the real one\"… or is this an approximate calculation?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1046609,
      "author_name": "",
      "author_url": "",
      "post_date": "10/11/2020 20:40:28",
      "content": "<p><a href=\"https://www.kaggle.com/sapr3s\" target=\"_blank\">@sapr3s</a>  i made a mistake in my original post, it's more like 10ms per call instead of 100, to measure it i just made a few measurments by putting a %%time decorator while executing the API call :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1047925,
      "author_name": "abhimanyud",
      "author_url": "",
      "post_date": "10/13/2020 04:14:12",
      "content": "<p>This is just perfect!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1046593": "As this competition pose a great dataflow optimization challenge, i wrote a kernel to generate a fake test set from the train set in a similar fashion as the real api from the competition, for debugging and optimisation purpose.\n\nA few differences to observe compared to the real api:\n- the iterator need an index to be called\n- you don't haveto call the function predict to go to the next batch\n- the call is a bit faster than the real api, so you might want to add a delay to simulate a real call (my api : 1ms / call against about 10ms per call for the real one)\n\n\nThe kernel can be found here : https://www.kaggle.com/rously/fake-api-generation-notebook-riid\nA preprocessed test set of about 2.4M lines can be found here:  https://www.kaggle.com/rously/fake-api-generation-riid-preprocessed\n\nIf you have any improvemement idea for this kernel, please feel free to leave in comments",
    "1046607": "Thank you, I was just planning on doing something similar. How you measured \"100ms per call for the real one\"... or is this an approximate calculation?",
    "1046609": "sapr3s  i made a mistake in my original post, it's more like 10ms per call instead of 100, to measure it i just made a few measurments by putting a %%time decorator while executing the API call :)",
    "1047925": "This is just perfect!"
  },
  "source": "meta"
}