{
  "id": 209598,
  "title": "Incompatibilities between the Kaggle platform and this kind of dataset",
  "url": "/competitions/riiid-test-answer-prediction/discussion/209598",
  "author_name": "",
  "post_date": "2021-01-08T01:17:48.203760100Z",
  "votes": 24,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Terribly sorry, I <a href=\"https://en.wikiquote.org/wiki/Blaise_Pascal#Pascal_plus_longue\" target=\"_blank\">appear to have written a book</a>.</p>\n<h1>Intro/context</h1>\n<p>This was my first Kaggle competition - in fact, I joined Kaggle specifically because of it, attracted by the promise of an interesting educational dataset.</p>\n<p>Overall, I am actually very impressed, and often pleasantly surprised, by Kaggle as a platform and a community. It feels like a very nice place to be.</p>\n<p>But on the other hand, this will probably also be my last <em>code</em> competition on Kaggle for a good long while, because of the amount of toil (as distinct from work) and confusion that it involved. I estimate that around <strong>80-90%</strong> of the time I spent on this competition was spent on things that decidedly are <strong>not</strong> data science, machine learning, data analysis, or anything that drew my interest initially. Most of that time was spent wrestling with the structure and size of the data, and the many … <em>surprising</em> ways that they interacted with the Kaggle platform. It seems to me (or at least I hope) that these interactions were not at all obvious to the designers of the challenge and/or platform beforehand, so that's why I feel compelled to point them out.</p>\n<p>I have seen people on Kaggle use the phrase \"Data Scientists are not Software Engineers\" to try and moderate expectations of what a median Kaggler ought to be able to do in terms of programming, software, knowing what the computer is doing \"under the hood\", etc. This, I think, is very reasonable - the more that you are forced to focus on the implementation details, the less time and mental energy you have to spend on actually doing and/or learning data science.</p>\n<p>That said, I <strong>am</strong> a (former) software engineer. With 10 years of professional experience. 6 of them at Google. If software engineering required this kind of toil, I would have quit a <em>lot</em> earlier. I say this not to brag or complain (necessarily). And also not to imply that Kaggle should just do things \"the Google Way\" (in fact, many of the things I love about Kaggle so far would be probably impossible if it was operating as a Google Product). I just want to hopefully give appropriate context to how much of a problem these problems are.</p>\n<h1>Analyzing Data</h1>\n<p>The main part of the Riiid dataset (the training data itself) is contained a singe .csv file that is around 5GB in size. Some aspects of the data also have a somewhat tricky structure, including:</p>\n<ul>\n<li>a \"content_id\" field that does not, in fact, uniquely identify content. (instead, it's a combination of two unrelated IDs <strong>which intersect in values</strong>). This means that operations like identifying a piece of content, or filtering for types of content, are non-trivial, and (even when implemented correctly <em>and</em> efficiently) can frequently consume more processor time and memory than an actual id.</li>\n<li>two fields (prior_question_had_explanation, prior_question_elapsed_time) whose values relate to <strong>some other</strong> row in the data. Or, actually <strong>some set of other rows</strong> in some cases. Because \"prior_question\" is actually somewhat a misnomer, since these values are associated with \"bundles\" of questions which may or may not contain exactly one question.</li>\n<li>This is kind of, but not really, time series data - it's really many different independent (or forced-independent) time series, one for each user of the app, forced together by the input file and output API.</li>\n</ul>\n<p>Whether these properties of the dataset are, by themselves, a bad idea is out of the scope of this particular post. Instead, like I said, my focus here is how they <em>interact</em> with Kaggle as a platform.</p>\n<p>The format of this data challenge - and, from what I've seen, Kaggle in general - leans heavily on using Pandas to process data, and storing state in .csv files:</p>\n<ul>\n<li>the tutorial/API overview focused exclusively on loading the data with Pandas</li>\n<li>the API itself serves the chunks of test data in Pandas dataframes, and there are no other options offered.</li>\n</ul>\n<p>I imagine that normally, this has the purpose of lowering the barrier to entry by limiting the number of tools that a person has to know to get started. But in this case, I think it had the complete opposite effect - the consensus seems to be that it's impossible, or at the very least unwise, to try and deal with this dataset using Pandas alone. So instead, everybody had to go and figure out, individually, some set of tools that actually works for their purpose. Because the starter code and explanation provided for the competition did not adequately start people on the path to dealing with this data successfully.</p>\n<h3>Memory</h3>\n<p>The main pain point which arises from the interaction between this dataset, Pandas, and the Kaggle platform is this:</p>\n<ul>\n<li>Pandas tends to be very memory-hungry, often apparently making multiple intermediate copies of the data it's working with.<ul>\n<li>it's not always easy to predict when and why it will do this</li></ul></li>\n<li>Kaggle notebooks are limited to 16GB of RAM, which is roughly 3x the size of the training data</li>\n<li>When a notebook runs out of those 16GB, <strong>it stops working or restarts, losing intermediate state</strong></li>\n<li>What's even worse, when a notebook runs out of memory in non-interactive mode (i.e. on submit), <strong>it hangs until it runs out of time</strong> instead of stopping and reporting the problem.<ul>\n<li>Apparently, this has been a \"<a href=\"https://www.kaggle.com/product-feedback/71176\" target=\"_blank\">known issue, will not fix</a>\" for at least two years. The advice is basically to test it in interactive mode first. But this isn't very practical if, say, the memory failure only happens 5 hours into the processing - which is quite likely and (for me) common when trying to deal with this dataset.<ul>\n<li>Surely, if it's possible to <em>know</em> I'm out of memory in interactive mode, it should be possible to at least <em>guess</em> this is happening in the other mode, and alert me to this fact, so that I don't have to interactively watch my batch submission to manually guess whether it has actually failed 3 hours ago?</li></ul></li>\n<li>This issue alone - which is undocumented, except for that two year old forum post - cost me a couple of days <strong>just trying to understand what is happening</strong></li>\n<li>Thrashing memory for 9 hours because it ran out of memory in the  first 10 minutes can't possibly be a good use of resources?..</li></ul></li>\n</ul>\n<p>Of course, this problem isn't limited to Pandas. Fundamentally, 5GB of data is sitting in a single file, and it needs to go into memory, and get manipulated, all without ever going over 16GB total in memory, <strong>and</strong> without going over 9 hours of processing time for whatever operation you're trying to do.</p>\n<p>This problem is exacerbated when complex operations are, in fact, necessary to get information out of the data which is present but misaligned, as with the prior_question fields. In the end I did manage to associate these fields with the appropriate row in the table, but I had to find a very precise way of shuffling the data between Pandas and Datatable to avoid their respective weak points.</p>\n<p>If these rows had not been shifted in the first place - if the training data had this_question_had_explanation and this_question_elapsed_time - I think many more people would have been much less confused, and more able to come up with interesting ways of using that data. Just search for \"prior_question\" in the discussion forums and marvel at the length and confusion of the posts.</p>\n<p>I can only imagine that the motivation behind shifting the data was to make it more clear that these particular fields couldn't be used directly in training, but again, I think it created way more confusion than it resolved. I think any number of alternative solutions would have worked so much better:</p>\n<ul>\n<li>Not shifting, and just treating those fields the same exact way as answered_correctly in the test data  (and providing a clear example of using this test data)</li>\n<li>shifting, but providing an additional field(or fields) which tracks what the prior_question <strong>was</strong></li>\n<li>shifting, but having a less complicated and better-explained scheme about what \"prior\" means<ul>\n<li>Actually have people on hand to explain this in the forums, instead of leaving us to guess and experiment</li></ul></li>\n</ul>\n<h3>Notebooks</h3>\n<p>This is a much smaller issue than the memory/data size/data format interaction. But I often found the options for managing, versioning and chaining notebooks to be quite limiting.</p>\n<p>I ended up making 40+ notebooks for this data challenge. Some are just exploratory, some are chained together in that one notebook produces output which another notebook then uses. I think that if the data was more manageable, the number and variety of notebooks would probably also go down and be more manageable. But as is, I needed to keep up with many different independent bits, all of which had several different (potentially broken) versions.</p>\n<p>Specific things I found lacking:</p>\n<ul>\n<li>If I have to find a particular notebook, my only option is to stare at a flat list of names and hope I remember what I named it. This is exacerbated by the quite-short character limit on notebook names.</li>\n<li>There is clearly a version control system in place, but frustratingly few capabilities are exposed to us. In particular, I found myself wanting (but unable) to:<ul>\n<li>revert to a previous version, or even just discard uncommitted changes<ul>\n<li>or heck, even just diff my uncommitted changes with the committed version</li></ul></li>\n<li>use a specific previous version's output as input (e.g. either because the current one is broken - and will be for another 5 hours as the fix re-runs)</li>\n<li>fork from a previous version (or even just from the committed current version I'm looking at when I press \"fork\", instead of the draft)</li></ul></li>\n<li>I also would very much appreciate some kind of graph view of my related notebooks: which ones are forked from which? which ones feed data into which? I honestly couldn't tell you right now what my 40+ notebooks do and how that happened. but if I had a graph, I might.</li>\n</ul>\n<h1>Submission process</h1>\n<p>It took me 12 submissions over 3 days just to get by baseline code to submit, and after that it took me about 10 more to add just one more bit of logic to it. After that, I ran out of time, so I really didn't get to do a lot of interesting stuff. Then again, running out of time might have better for my sanity than the alternative.</p>\n<h3>Debugging and API shape</h3>\n<p>I think lots of people have already talked about this extensively, so I'll try to keep at least this section brief:</p>\n<ul>\n<li>The fact that there is no information about why a submission failed is absolutely unworkable<ul>\n<li>It's not even easy to tell <strong>how fast</strong> my submission failed unless I am watching the submission run (and then also recording it myself somewhere) - e.g. did this submission from a day ago fail after 5 minutes because I have a bug, or after 9 hours because it timed out? </li></ul></li>\n<li>This could have been mitigated by a robust test set which includes all known edge cases. Instead, the set of data that's visible when we are testing the API seems to have exactly <strong>one</strong> edge case (new user), and is conspicuously missing another edge case which the organizers already knew was important (lectures) - they put the (one-sentence) description of this edge case in bold in their starter notebook.<ul>\n<li>Possibly even a small synthetic data set which covers lots of edge cases would have been a major improvement<ul>\n<li>of course, these edge cases would have to be found first before the competition starts. Perhaps through internal testing?..</li></ul></li></ul></li>\n<li>This is all further  exacerbated by the fact that the API shape has several gotchas.<ul>\n<li>The main one is that the predictions and the subsequent \"correct answers to your previous submission\" have a different shape: one has -1 for lectures, the other one MUST skip lectures. This was not explicitly explained by organizers, and again, not testable.</li>\n<li>The part where the correct answers are a serialized bit of python shoved into some arbitrary single field in an unrelated dataframe… Actually was far less problematic than I thought it would be. But it <em>is</em> a strong indicator that this API was not designed for this challenge - that somebody took a well-thought-out API for a different kind of problem (actual time series data, without a need to track state of previous submissions) and tried to adapt it by making the smallest possible number of changes. Instead of designing an API to specifically fit this dataset.<ul>\n<li>I can see how this decision, too, may have been made in the name of simplicity (people are already used to this API…), but it only made things more confusing and complicated.</li></ul></li></ul></li>\n<li>In these circumstances, the 5 submission per day limit is extremely frustrating. For that matter, why do failed submissions <strong>ever</strong> count toward that limit?</li>\n</ul>\n<h3>State</h3>\n<p>Ohhh boy, and I thought the prior_question fields were problematic when I was trying to process training data.</p>\n<p>This challenge is - or ought to be - a natural fit for tracking user state. Because we are trying to predict user actions based on previous user actions (well, and the app's choices based on user actions, but that's a topic for a whole different post). However, <strong>previous user state is split up in such a way that it's extremely hard to put it back together</strong>.</p>\n<p>We have some data (answered_correctly, content_id, …) associated with <strong>the previous chunk of test data</strong> (or, <strong>just the first time</strong>, from the state at the end of training). And some data (prior_question fields) associated with <strong>that user's previous most recent answer</strong>, which may come from <em>some</em> other chunk of test data, or the training data, <strong>at any point in the test process</strong>.</p>\n<p>So when I'm trying to actually collate all this and understand, for example, <strong>whether a user answered their previous question correctly, how long it took them to answer that question, what that question was, and whether they received feedback</strong>, how do I do that?.. Suppose I am tracking general \"last known user state\" and updating it as I get information from the various sources.</p>\n<ul>\n<li><strong>At what point does the prior_question data actually line up with the answered_correctly data, so they both refer to the same question?</strong></li>\n<li>Suppose I also care about question counts, and dividing various totals by the number of questions the user has seen. Can I guarantee that just one question count is actually up-to-date for all the fields simultaneously?..</li>\n</ul>\n<p>This might actually make for an interesting Software Engineer interview question. But in simplified form, on a whiteboard, and without the complete lack of feedback. </p>\n<p>I know for a fact that I have at least one bug in my logic for dealing with this. But whenever I touched that code, my entire submission would start failing again. And I just did not have enough submission budget to experiment with it.</p>\n<h3>Time</h3>\n<p>We are given 9 hours minus 15 minutes (according to the instructions, the API takes up 15 minutes for itself) to predict 2.5 million data points. That means about 80-100 predictions a second (depending on load/spin-up time of my own data). This is motivated by saying that these predictions should be usable in a real-time setting. But no real-time setting needs predictions to be made in 0.01 of a second, while keeping the context of the entire userbase in RAM. Consider that a reasonable screen refresh rate is 30 frames a second. Making predictions at 3x that speed is pretty useless optimization. The reality, of course, is that 9 hours is the standard Kaggle notebook running time, and it couldn't be changed for this competition. In that case, though, the size of the test dataset should have been changed.</p>\n<p>Of course, making 100 predicions a second is only a problem when you want to keep complex user state and update it with elaborate logic. </p>\n<p>I guess you could argue that severely limiting the feasibility of tracking user state actually \"simplifies\" the challenge somewhat by encouraging <em>everyone</em> to not track complex state manually; instead, either rely on your model to implicitly track an approximation of it, or just throw out annoying state altogether. But this kind of \"simplification\" is sort of like simplifying <a href=\"https://en.wikipedia.org/wiki/Streetlight_effect\" target=\"_blank\">the process of searching for your keys in the dark</a> by smashing some of the streetlights, in order to make the feasible search space smaller.</p>\n<h1>Conclusions</h1>\n<p>The ideal (given unlimited resources) solution to situations like this, I think, would be to:</p>\n<ul>\n<li>Store data like this in a database at the outset, not a text file</li>\n<li>Provide examples of using appropriate tools for the data format and optimization goal (i.e. for predicting correctness in small chunks - probably not (just) pandas)</li>\n<li>Don't structure the data in a way that requires either (1) esoteric processing or (2) effectively ignoring most of the usefulness of that part of the data</li>\n<li>Design the (size and structure) of the evaluation part of the competition to be reasonably compatible with Kaggle's performance constraints</li>\n<li>Design the submission API to fit the specific dataset</li>\n<li>Provide robust tools and/or data sets for diagnosing and debugging problems <ul>\n<li>Maybe something like the <a href=\"https://www.kaggle.com/its7171/time-series-api-iter-test-emulator\" target=\"_blank\">emulator</a> made by one of the community members, but actually guaranteed to match the API</li></ul></li>\n<li><a href=\"https://en.wikipedia.org/wiki/Eating_your_own_dog_food\" target=\"_blank\">Dogfood</a> and make adjustments when it's still easy/possible to change the format of the challenge<ul>\n<li>And I don't mean \"do a run the pre-existing model around which this competition was designed\". I mean get other people, who were not involved in designing the challenge, to earnestly try to do some interesting stuff with the data and go through the submission flow on their own.</li></ul></li>\n</ul>",
  "messages": [
    {
      "id": "1143621",
      "postDate": "01/08/2021 01:17:48",
      "content": "<p>Terribly sorry, I <a href=\"https://en.wikiquote.org/wiki/Blaise_Pascal#Pascal_plus_longue\" target=\"_blank\">appear to have written a book</a>.</p>\n<h1>Intro/context</h1>\n<p>This was my first Kaggle competition - in fact, I joined Kaggle specifically because of it, attracted by the promise of an interesting educational dataset.</p>\n<p>Overall, I am actually very impressed, and often pleasantly surprised, by Kaggle as a platform and a community. It feels like a very nice place to be.</p>\n<p>But on the other hand, this will probably also be my last <em>code</em> competition on Kaggle for a good long while, because of the amount of toil (as distinct from work) and confusion that it involved. I estimate that around <strong>80-90%</strong> of the time I spent on this competition was spent on things that decidedly are <strong>not</strong> data science, machine learning, data analysis, or anything that drew my interest initially. Most of that time was spent wrestling with the structure and size of the data, and the many … <em>surprising</em> ways that they interacted with the Kaggle platform. It seems to me (or at least I hope) that these interactions were not at all obvious to the designers of the challenge and/or platform beforehand, so that's why I feel compelled to point them out.</p>\n<p>I have seen people on Kaggle use the phrase \"Data Scientists are not Software Engineers\" to try and moderate expectations of what a median Kaggler ought to be able to do in terms of programming, software, knowing what the computer is doing \"under the hood\", etc. This, I think, is very reasonable - the more that you are forced to focus on the implementation details, the less time and mental energy you have to spend on actually doing and/or learning data science.</p>\n<p>That said, I <strong>am</strong> a (former) software engineer. With 10 years of professional experience. 6 of them at Google. If software engineering required this kind of toil, I would have quit a <em>lot</em> earlier. I say this not to brag or complain (necessarily). And also not to imply that Kaggle should just do things \"the Google Way\" (in fact, many of the things I love about Kaggle so far would be probably impossible if it was operating as a Google Product). I just want to hopefully give appropriate context to how much of a problem these problems are.</p>\n<h1>Analyzing Data</h1>\n<p>The main part of the Riiid dataset (the training data itself) is contained a singe .csv file that is around 5GB in size. Some aspects of the data also have a somewhat tricky structure, including:</p>\n<ul>\n<li>a \"content_id\" field that does not, in fact, uniquely identify content. (instead, it's a combination of two unrelated IDs <strong>which intersect in values</strong>). This means that operations like identifying a piece of content, or filtering for types of content, are non-trivial, and (even when implemented correctly <em>and</em> efficiently) can frequently consume more processor time and memory than an actual id.</li>\n<li>two fields (prior_question_had_explanation, prior_question_elapsed_time) whose values relate to <strong>some other</strong> row in the data. Or, actually <strong>some set of other rows</strong> in some cases. Because \"prior_question\" is actually somewhat a misnomer, since these values are associated with \"bundles\" of questions which may or may not contain exactly one question.</li>\n<li>This is kind of, but not really, time series data - it's really many different independent (or forced-independent) time series, one for each user of the app, forced together by the input file and output API.</li>\n</ul>\n<p>Whether these properties of the dataset are, by themselves, a bad idea is out of the scope of this particular post. Instead, like I said, my focus here is how they <em>interact</em> with Kaggle as a platform.</p>\n<p>The format of this data challenge - and, from what I've seen, Kaggle in general - leans heavily on using Pandas to process data, and storing state in .csv files:</p>\n<ul>\n<li>the tutorial/API overview focused exclusively on loading the data with Pandas</li>\n<li>the API itself serves the chunks of test data in Pandas dataframes, and there are no other options offered.</li>\n</ul>\n<p>I imagine that normally, this has the purpose of lowering the barrier to entry by limiting the number of tools that a person has to know to get started. But in this case, I think it had the complete opposite effect - the consensus seems to be that it's impossible, or at the very least unwise, to try and deal with this dataset using Pandas alone. So instead, everybody had to go and figure out, individually, some set of tools that actually works for their purpose. Because the starter code and explanation provided for the competition did not adequately start people on the path to dealing with this data successfully.</p>\n<h3>Memory</h3>\n<p>The main pain point which arises from the interaction between this dataset, Pandas, and the Kaggle platform is this:</p>\n<ul>\n<li>Pandas tends to be very memory-hungry, often apparently making multiple intermediate copies of the data it's working with.<ul>\n<li>it's not always easy to predict when and why it will do this</li></ul></li>\n<li>Kaggle notebooks are limited to 16GB of RAM, which is roughly 3x the size of the training data</li>\n<li>When a notebook runs out of those 16GB, <strong>it stops working or restarts, losing intermediate state</strong></li>\n<li>What's even worse, when a notebook runs out of memory in non-interactive mode (i.e. on submit), <strong>it hangs until it runs out of time</strong> instead of stopping and reporting the problem.<ul>\n<li>Apparently, this has been a \"<a href=\"https://www.kaggle.com/product-feedback/71176\" target=\"_blank\">known issue, will not fix</a>\" for at least two years. The advice is basically to test it in interactive mode first. But this isn't very practical if, say, the memory failure only happens 5 hours into the processing - which is quite likely and (for me) common when trying to deal with this dataset.<ul>\n<li>Surely, if it's possible to <em>know</em> I'm out of memory in interactive mode, it should be possible to at least <em>guess</em> this is happening in the other mode, and alert me to this fact, so that I don't have to interactively watch my batch submission to manually guess whether it has actually failed 3 hours ago?</li></ul></li>\n<li>This issue alone - which is undocumented, except for that two year old forum post - cost me a couple of days <strong>just trying to understand what is happening</strong></li>\n<li>Thrashing memory for 9 hours because it ran out of memory in the  first 10 minutes can't possibly be a good use of resources?..</li></ul></li>\n</ul>\n<p>Of course, this problem isn't limited to Pandas. Fundamentally, 5GB of data is sitting in a single file, and it needs to go into memory, and get manipulated, all without ever going over 16GB total in memory, <strong>and</strong> without going over 9 hours of processing time for whatever operation you're trying to do.</p>\n<p>This problem is exacerbated when complex operations are, in fact, necessary to get information out of the data which is present but misaligned, as with the prior_question fields. In the end I did manage to associate these fields with the appropriate row in the table, but I had to find a very precise way of shuffling the data between Pandas and Datatable to avoid their respective weak points.</p>\n<p>If these rows had not been shifted in the first place - if the training data had this_question_had_explanation and this_question_elapsed_time - I think many more people would have been much less confused, and more able to come up with interesting ways of using that data. Just search for \"prior_question\" in the discussion forums and marvel at the length and confusion of the posts.</p>\n<p>I can only imagine that the motivation behind shifting the data was to make it more clear that these particular fields couldn't be used directly in training, but again, I think it created way more confusion than it resolved. I think any number of alternative solutions would have worked so much better:</p>\n<ul>\n<li>Not shifting, and just treating those fields the same exact way as answered_correctly in the test data  (and providing a clear example of using this test data)</li>\n<li>shifting, but providing an additional field(or fields) which tracks what the prior_question <strong>was</strong></li>\n<li>shifting, but having a less complicated and better-explained scheme about what \"prior\" means<ul>\n<li>Actually have people on hand to explain this in the forums, instead of leaving us to guess and experiment</li></ul></li>\n</ul>\n<h3>Notebooks</h3>\n<p>This is a much smaller issue than the memory/data size/data format interaction. But I often found the options for managing, versioning and chaining notebooks to be quite limiting.</p>\n<p>I ended up making 40+ notebooks for this data challenge. Some are just exploratory, some are chained together in that one notebook produces output which another notebook then uses. I think that if the data was more manageable, the number and variety of notebooks would probably also go down and be more manageable. But as is, I needed to keep up with many different independent bits, all of which had several different (potentially broken) versions.</p>\n<p>Specific things I found lacking:</p>\n<ul>\n<li>If I have to find a particular notebook, my only option is to stare at a flat list of names and hope I remember what I named it. This is exacerbated by the quite-short character limit on notebook names.</li>\n<li>There is clearly a version control system in place, but frustratingly few capabilities are exposed to us. In particular, I found myself wanting (but unable) to:<ul>\n<li>revert to a previous version, or even just discard uncommitted changes<ul>\n<li>or heck, even just diff my uncommitted changes with the committed version</li></ul></li>\n<li>use a specific previous version's output as input (e.g. either because the current one is broken - and will be for another 5 hours as the fix re-runs)</li>\n<li>fork from a previous version (or even just from the committed current version I'm looking at when I press \"fork\", instead of the draft)</li></ul></li>\n<li>I also would very much appreciate some kind of graph view of my related notebooks: which ones are forked from which? which ones feed data into which? I honestly couldn't tell you right now what my 40+ notebooks do and how that happened. but if I had a graph, I might.</li>\n</ul>\n<h1>Submission process</h1>\n<p>It took me 12 submissions over 3 days just to get by baseline code to submit, and after that it took me about 10 more to add just one more bit of logic to it. After that, I ran out of time, so I really didn't get to do a lot of interesting stuff. Then again, running out of time might have better for my sanity than the alternative.</p>\n<h3>Debugging and API shape</h3>\n<p>I think lots of people have already talked about this extensively, so I'll try to keep at least this section brief:</p>\n<ul>\n<li>The fact that there is no information about why a submission failed is absolutely unworkable<ul>\n<li>It's not even easy to tell <strong>how fast</strong> my submission failed unless I am watching the submission run (and then also recording it myself somewhere) - e.g. did this submission from a day ago fail after 5 minutes because I have a bug, or after 9 hours because it timed out? </li></ul></li>\n<li>This could have been mitigated by a robust test set which includes all known edge cases. Instead, the set of data that's visible when we are testing the API seems to have exactly <strong>one</strong> edge case (new user), and is conspicuously missing another edge case which the organizers already knew was important (lectures) - they put the (one-sentence) description of this edge case in bold in their starter notebook.<ul>\n<li>Possibly even a small synthetic data set which covers lots of edge cases would have been a major improvement<ul>\n<li>of course, these edge cases would have to be found first before the competition starts. Perhaps through internal testing?..</li></ul></li></ul></li>\n<li>This is all further  exacerbated by the fact that the API shape has several gotchas.<ul>\n<li>The main one is that the predictions and the subsequent \"correct answers to your previous submission\" have a different shape: one has -1 for lectures, the other one MUST skip lectures. This was not explicitly explained by organizers, and again, not testable.</li>\n<li>The part where the correct answers are a serialized bit of python shoved into some arbitrary single field in an unrelated dataframe… Actually was far less problematic than I thought it would be. But it <em>is</em> a strong indicator that this API was not designed for this challenge - that somebody took a well-thought-out API for a different kind of problem (actual time series data, without a need to track state of previous submissions) and tried to adapt it by making the smallest possible number of changes. Instead of designing an API to specifically fit this dataset.<ul>\n<li>I can see how this decision, too, may have been made in the name of simplicity (people are already used to this API…), but it only made things more confusing and complicated.</li></ul></li></ul></li>\n<li>In these circumstances, the 5 submission per day limit is extremely frustrating. For that matter, why do failed submissions <strong>ever</strong> count toward that limit?</li>\n</ul>\n<h3>State</h3>\n<p>Ohhh boy, and I thought the prior_question fields were problematic when I was trying to process training data.</p>\n<p>This challenge is - or ought to be - a natural fit for tracking user state. Because we are trying to predict user actions based on previous user actions (well, and the app's choices based on user actions, but that's a topic for a whole different post). However, <strong>previous user state is split up in such a way that it's extremely hard to put it back together</strong>.</p>\n<p>We have some data (answered_correctly, content_id, …) associated with <strong>the previous chunk of test data</strong> (or, <strong>just the first time</strong>, from the state at the end of training). And some data (prior_question fields) associated with <strong>that user's previous most recent answer</strong>, which may come from <em>some</em> other chunk of test data, or the training data, <strong>at any point in the test process</strong>.</p>\n<p>So when I'm trying to actually collate all this and understand, for example, <strong>whether a user answered their previous question correctly, how long it took them to answer that question, what that question was, and whether they received feedback</strong>, how do I do that?.. Suppose I am tracking general \"last known user state\" and updating it as I get information from the various sources.</p>\n<ul>\n<li><strong>At what point does the prior_question data actually line up with the answered_correctly data, so they both refer to the same question?</strong></li>\n<li>Suppose I also care about question counts, and dividing various totals by the number of questions the user has seen. Can I guarantee that just one question count is actually up-to-date for all the fields simultaneously?..</li>\n</ul>\n<p>This might actually make for an interesting Software Engineer interview question. But in simplified form, on a whiteboard, and without the complete lack of feedback. </p>\n<p>I know for a fact that I have at least one bug in my logic for dealing with this. But whenever I touched that code, my entire submission would start failing again. And I just did not have enough submission budget to experiment with it.</p>\n<h3>Time</h3>\n<p>We are given 9 hours minus 15 minutes (according to the instructions, the API takes up 15 minutes for itself) to predict 2.5 million data points. That means about 80-100 predictions a second (depending on load/spin-up time of my own data). This is motivated by saying that these predictions should be usable in a real-time setting. But no real-time setting needs predictions to be made in 0.01 of a second, while keeping the context of the entire userbase in RAM. Consider that a reasonable screen refresh rate is 30 frames a second. Making predictions at 3x that speed is pretty useless optimization. The reality, of course, is that 9 hours is the standard Kaggle notebook running time, and it couldn't be changed for this competition. In that case, though, the size of the test dataset should have been changed.</p>\n<p>Of course, making 100 predicions a second is only a problem when you want to keep complex user state and update it with elaborate logic. </p>\n<p>I guess you could argue that severely limiting the feasibility of tracking user state actually \"simplifies\" the challenge somewhat by encouraging <em>everyone</em> to not track complex state manually; instead, either rely on your model to implicitly track an approximation of it, or just throw out annoying state altogether. But this kind of \"simplification\" is sort of like simplifying <a href=\"https://en.wikipedia.org/wiki/Streetlight_effect\" target=\"_blank\">the process of searching for your keys in the dark</a> by smashing some of the streetlights, in order to make the feasible search space smaller.</p>\n<h1>Conclusions</h1>\n<p>The ideal (given unlimited resources) solution to situations like this, I think, would be to:</p>\n<ul>\n<li>Store data like this in a database at the outset, not a text file</li>\n<li>Provide examples of using appropriate tools for the data format and optimization goal (i.e. for predicting correctness in small chunks - probably not (just) pandas)</li>\n<li>Don't structure the data in a way that requires either (1) esoteric processing or (2) effectively ignoring most of the usefulness of that part of the data</li>\n<li>Design the (size and structure) of the evaluation part of the competition to be reasonably compatible with Kaggle's performance constraints</li>\n<li>Design the submission API to fit the specific dataset</li>\n<li>Provide robust tools and/or data sets for diagnosing and debugging problems <ul>\n<li>Maybe something like the <a href=\"https://www.kaggle.com/its7171/time-series-api-iter-test-emulator\" target=\"_blank\">emulator</a> made by one of the community members, but actually guaranteed to match the API</li></ul></li>\n<li><a href=\"https://en.wikipedia.org/wiki/Eating_your_own_dog_food\" target=\"_blank\">Dogfood</a> and make adjustments when it's still easy/possible to change the format of the challenge<ul>\n<li>And I don't mean \"do a run the pre-existing model around which this competition was designed\". I mean get other people, who were not involved in designing the challenge, to earnestly try to do some interesting stuff with the data and go through the submission flow on their own.</li></ul></li>\n</ul>",
      "rawMarkdown": "Terribly sorry, I [appear to have written a book](https://en.wikiquote.org/wiki/Blaise_Pascal#Pascal_plus_longue).\n\n# Intro/context\n\nThis was my first Kaggle competition - in fact, I joined Kaggle specifically because of it, attracted by the promise of an interesting educational dataset.\n\nOverall, I am actually very impressed, and often pleasantly surprised, by Kaggle as a platform and a community. It feels like a very nice place to be.\n\nBut on the other hand, this will probably also be my last *code* competition on Kaggle for a good long while, because of the amount of toil (as distinct from work) and confusion that it involved. I estimate that around **80-90%** of the time I spent on this competition was spent on things that decidedly are **not** data science, machine learning, data analysis, or anything that drew my interest initially. Most of that time was spent wrestling with the structure and size of the data, and the many ... *surprising* ways that they interacted with the Kaggle platform. It seems to me (or at least I hope) that these interactions were not at all obvious to the designers of the challenge and/or platform beforehand, so that's why I feel compelled to point them out.\n\nI have seen people on Kaggle use the phrase \"Data Scientists are not Software Engineers\" to try and moderate expectations of what a median Kaggler ought to be able to do in terms of programming, software, knowing what the computer is doing \"under the hood\", etc. This, I think, is very reasonable - the more that you are forced to focus on the implementation details, the less time and mental energy you have to spend on actually doing and/or learning data science.\n\nThat said, I **am** a (former) software engineer. With 10 years of professional experience. 6 of them at Google. If software engineering required this kind of toil, I would have quit a *lot* earlier. I say this not to brag or complain (necessarily). And also not to imply that Kaggle should just do things \"the Google Way\" (in fact, many of the things I love about Kaggle so far would be probably impossible if it was operating as a Google Product). I just want to hopefully give appropriate context to how much of a problem these problems are.\n\n# Analyzing Data\n\nThe main part of the Riiid dataset (the training data itself) is contained a singe .csv file that is around 5GB in size. Some aspects of the data also have a somewhat tricky structure, including:\n\n- a \"content_id\" field that does not, in fact, uniquely identify content. (instead, it's a combination of two unrelated IDs **which intersect in values**). This means that operations like identifying a piece of content, or filtering for types of content, are non-trivial, and (even when implemented correctly *and* efficiently) can frequently consume more processor time and memory than an actual id.\n- two fields (prior_question_had_explanation, prior_question_elapsed_time) whose values relate to **some other** row in the data. Or, actually **some set of other rows** in some cases. Because \"prior_question\" is actually somewhat a misnomer, since these values are associated with \"bundles\" of questions which may or may not contain exactly one question.\n- This is kind of, but not really, time series data - it's really many different independent (or forced-independent) time series, one for each user of the app, forced together by the input file and output API.\n\nWhether these properties of the dataset are, by themselves, a bad idea is out of the scope of this particular post. Instead, like I said, my focus here is how they *interact* with Kaggle as a platform.\n\nThe format of this data challenge - and, from what I've seen, Kaggle in general - leans heavily on using Pandas to process data, and storing state in .csv files:\n\n- the tutorial/API overview focused exclusively on loading the data with Pandas\n- the API itself serves the chunks of test data in Pandas dataframes, and there are no other options offered.\n\nI imagine that normally, this has the purpose of lowering the barrier to entry by limiting the number of tools that a person has to know to get started. But in this case, I think it had the complete opposite effect - the consensus seems to be that it's impossible, or at the very least unwise, to try and deal with this dataset using Pandas alone. So instead, everybody had to go and figure out, individually, some set of tools that actually works for their purpose. Because the starter code and explanation provided for the competition did not adequately start people on the path to dealing with this data successfully.\n\n### Memory\nThe main pain point which arises from the interaction between this dataset, Pandas, and the Kaggle platform is this:\n\n- Pandas tends to be very memory-hungry, often apparently making multiple intermediate copies of the data it's working with.\n    - it's not always easy to predict when and why it will do this\n- Kaggle notebooks are limited to 16GB of RAM, which is roughly 3x the size of the training data\n- When a notebook runs out of those 16GB, **it stops working or restarts, losing intermediate state**\n- What's even worse, when a notebook runs out of memory in non-interactive mode (i.e. on submit), **it hangs until it runs out of time** instead of stopping and reporting the problem.\n    - Apparently, this has been a \"[known issue, will not fix](https://www.kaggle.com/product-feedback/71176)\" for at least two years. The advice is basically to test it in interactive mode first. But this isn't very practical if, say, the memory failure only happens 5 hours into the processing - which is quite likely and (for me) common when trying to deal with this dataset.\n        - Surely, if it's possible to *know* I'm out of memory in interactive mode, it should be possible to at least *guess* this is happening in the other mode, and alert me to this fact, so that I don't have to interactively watch my batch submission to manually guess whether it has actually failed 3 hours ago?\n    - This issue alone - which is undocumented, except for that two year old forum post - cost me a couple of days **just trying to understand what is happening**\n    - Thrashing memory for 9 hours because it ran out of memory in the  first 10 minutes can't possibly be a good use of resources?..\n\n    \nOf course, this problem isn't limited to Pandas. Fundamentally, 5GB of data is sitting in a single file, and it needs to go into memory, and get manipulated, all without ever going over 16GB total in memory, **and** without going over 9 hours of processing time for whatever operation you're trying to do.\n\nThis problem is exacerbated when complex operations are, in fact, necessary to get information out of the data which is present but misaligned, as with the prior_question fields. In the end I did manage to associate these fields with the appropriate row in the table, but I had to find a very precise way of shuffling the data between Pandas and Datatable to avoid their respective weak points.\n\nIf these rows had not been shifted in the first place - if the training data had this_question_had_explanation and this_question_elapsed_time - I think many more people would have been much less confused, and more able to come up with interesting ways of using that data. Just search for \"prior_question\" in the discussion forums and marvel at the length and confusion of the posts.\n\nI can only imagine that the motivation behind shifting the data was to make it more clear that these particular fields couldn't be used directly in training, but again, I think it created way more confusion than it resolved. I think any number of alternative solutions would have worked so much better:\n\n- Not shifting, and just treating those fields the same exact way as answered_correctly in the test data  (and providing a clear example of using this test data)\n- shifting, but providing an additional field(or fields) which tracks what the prior_question **was**\n- shifting, but having a less complicated and better-explained scheme about what \"prior\" means\n    - Actually have people on hand to explain this in the forums, instead of leaving us to guess and experiment\n\n### Notebooks\n\nThis is a much smaller issue than the memory/data size/data format interaction. But I often found the options for managing, versioning and chaining notebooks to be quite limiting.\n\nI ended up making 40+ notebooks for this data challenge. Some are just exploratory, some are chained together in that one notebook produces output which another notebook then uses. I think that if the data was more manageable, the number and variety of notebooks would probably also go down and be more manageable. But as is, I needed to keep up with many different independent bits, all of which had several different (potentially broken) versions.\n\nSpecific things I found lacking:\n\n- If I have to find a particular notebook, my only option is to stare at a flat list of names and hope I remember what I named it. This is exacerbated by the quite-short character limit on notebook names.\n- There is clearly a version control system in place, but frustratingly few capabilities are exposed to us. In particular, I found myself wanting (but unable) to:\n    - revert to a previous version, or even just discard uncommitted changes\n        - or heck, even just diff my uncommitted changes with the committed version\n    - use a specific previous version's output as input (e.g. either because the current one is broken - and will be for another 5 hours as the fix re-runs)\n    - fork from a previous version (or even just from the committed current version I'm looking at when I press \"fork\", instead of the draft)\n- I also would very much appreciate some kind of graph view of my related notebooks: which ones are forked from which? which ones feed data into which? I honestly couldn't tell you right now what my 40+ notebooks do and how that happened. but if I had a graph, I might.\n\n# Submission process\n\nIt took me 12 submissions over 3 days just to get by baseline code to submit, and after that it took me about 10 more to add just one more bit of logic to it. After that, I ran out of time, so I really didn't get to do a lot of interesting stuff. Then again, running out of time might have better for my sanity than the alternative.\n\n### Debugging and API shape\n\nI think lots of people have already talked about this extensively, so I'll try to keep at least this section brief:\n\n- The fact that there is no information about why a submission failed is absolutely unworkable\n    - It's not even easy to tell **how fast** my submission failed unless I am watching the submission run (and then also recording it myself somewhere) - e.g. did this submission from a day ago fail after 5 minutes because I have a bug, or after 9 hours because it timed out? \n- This could have been mitigated by a robust test set which includes all known edge cases. Instead, the set of data that's visible when we are testing the API seems to have exactly **one** edge case (new user), and is conspicuously missing another edge case which the organizers already knew was important (lectures) - they put the (one-sentence) description of this edge case in bold in their starter notebook.\n    - Possibly even a small synthetic data set which covers lots of edge cases would have been a major improvement\n        - of course, these edge cases would have to be found first before the competition starts. Perhaps through internal testing?..\n- This is all further  exacerbated by the fact that the API shape has several gotchas.\n    - The main one is that the predictions and the subsequent \"correct answers to your previous submission\" have a different shape: one has -1 for lectures, the other one MUST skip lectures. This was not explicitly explained by organizers, and again, not testable.\n    - The part where the correct answers are a serialized bit of python shoved into some arbitrary single field in an unrelated dataframe... Actually was far less problematic than I thought it would be. But it *is* a strong indicator that this API was not designed for this challenge - that somebody took a well-thought-out API for a different kind of problem (actual time series data, without a need to track state of previous submissions) and tried to adapt it by making the smallest possible number of changes. Instead of designing an API to specifically fit this dataset.\n        - I can see how this decision, too, may have been made in the name of simplicity (people are already used to this API...), but it only made things more confusing and complicated.\n- In these circumstances, the 5 submission per day limit is extremely frustrating. For that matter, why do failed submissions **ever** count toward that limit?\n\n### State\n\nOhhh boy, and I thought the prior_question fields were problematic when I was trying to process training data.\n\nThis challenge is - or ought to be - a natural fit for tracking user state. Because we are trying to predict user actions based on previous user actions (well, and the app's choices based on user actions, but that's a topic for a whole different post). However, **previous user state is split up in such a way that it's extremely hard to put it back together**.\n\nWe have some data (answered_correctly, content_id, ...) associated with **the previous chunk of test data** (or, **just the first time**, from the state at the end of training). And some data (prior_question fields) associated with **that user's previous most recent answer**, which may come from *some* other chunk of test data, or the training data, **at any point in the test process**.\n\nSo when I'm trying to actually collate all this and understand, for example, **whether a user answered their previous question correctly, how long it took them to answer that question, what that question was, and whether they received feedback**, how do I do that?.. Suppose I am tracking general \"last known user state\" and updating it as I get information from the various sources.\n\n- **At what point does the prior_question data actually line up with the answered_correctly data, so they both refer to the same question?**\n- Suppose I also care about question counts, and dividing various totals by the number of questions the user has seen. Can I guarantee that just one question count is actually up-to-date for all the fields simultaneously?..\n\nThis might actually make for an interesting Software Engineer interview question. But in simplified form, on a whiteboard, and without the complete lack of feedback. \n\nI know for a fact that I have at least one bug in my logic for dealing with this. But whenever I touched that code, my entire submission would start failing again. And I just did not have enough submission budget to experiment with it.\n\n### Time\n\nWe are given 9 hours minus 15 minutes (according to the instructions, the API takes up 15 minutes for itself) to predict 2.5 million data points. That means about 80-100 predictions a second (depending on load/spin-up time of my own data). This is motivated by saying that these predictions should be usable in a real-time setting. But no real-time setting needs predictions to be made in 0.01 of a second, while keeping the context of the entire userbase in RAM. Consider that a reasonable screen refresh rate is 30 frames a second. Making predictions at 3x that speed is pretty useless optimization. The reality, of course, is that 9 hours is the standard Kaggle notebook running time, and it couldn't be changed for this competition. In that case, though, the size of the test dataset should have been changed.\n\nOf course, making 100 predicions a second is only a problem when you want to keep complex user state and update it with elaborate logic. \n\nI guess you could argue that severely limiting the feasibility of tracking user state actually \"simplifies\" the challenge somewhat by encouraging *everyone* to not track complex state manually; instead, either rely on your model to implicitly track an approximation of it, or just throw out annoying state altogether. But this kind of \"simplification\" is sort of like simplifying [the process of searching for your keys in the dark](https://en.wikipedia.org/wiki/Streetlight_effect) by smashing some of the streetlights, in order to make the feasible search space smaller.\n\n# Conclusions\n\nThe ideal (given unlimited resources) solution to situations like this, I think, would be to:\n\n- Store data like this in a database at the outset, not a text file\n- Provide examples of using appropriate tools for the data format and optimization goal (i.e. for predicting correctness in small chunks - probably not (just) pandas)\n- Don't structure the data in a way that requires either (1) esoteric processing or (2) effectively ignoring most of the usefulness of that part of the data\n- Design the (size and structure) of the evaluation part of the competition to be reasonably compatible with Kaggle's performance constraints\n- Design the submission API to fit the specific dataset\n- Provide robust tools and/or data sets for diagnosing and debugging problems \n    - Maybe something like the [emulator](https://www.kaggle.com/its7171/time-series-api-iter-test-emulator) made by one of the community members, but actually guaranteed to match the API\n- [Dogfood](https://en.wikipedia.org/wiki/Eating_your_own_dog_food) and make adjustments when it's still easy/possible to change the format of the challenge\n    - And I don't mean \"do a run the pre-existing model around which this competition was designed\". I mean get other people, who were not involved in designing the challenge, to earnestly try to do some interesting stuff with the data and go through the submission flow on their own.",
      "votes": null
    },
    {
      "id": "1143865",
      "postDate": "01/08/2021 05:34:20",
      "content": "<p>I have read your topic with a full of sympathy (as many participants do).<br>\nI have been struggling with a lot of submission timeout, memory shortage, and other unknown errors.<br>\nKaggle should upgrade their system with more computational power and more user-friendly way.<br>\nHowever, from the positive side, I could learn how to pre-process a bit dirty time-series data, how to speed up processing,  how to handle data memory efficiently etc. This competition gives me a chance to learn a lot of things, which experienced engineers like you have already acquired.</p>",
      "rawMarkdown": "I have read your topic with a full of sympathy (as many participants do).\nI have been struggling with a lot of submission timeout, memory shortage, and other unknown errors.\nKaggle should upgrade their system with more computational power and more user-friendly way.\nHowever, from the positive side, I could learn how to pre-process a bit dirty time-series data, how to speed up processing,  how to handle data memory efficiently etc. This competition gives me a chance to learn a lot of things, which experienced engineers like you have already acquired.",
      "votes": null
    },
    {
      "id": "1144550",
      "postDate": "01/08/2021 14:25:19",
      "content": "<p>We all struggled to fit into the constraints of the API and overcome the difficulties of this data. However, people have demonstrated that it was possible to achieve amazing results using a small chunk of the data and/or smart feature pre-processing. Maybe all of this using only Kaggle notebooks.</p>\n<p>I understand that sometimes frustration can take over, it happened to me in past competitions. But please keep in mind that Kaggle competitions are not perfect, Kaggle notebooks are not perfect, data is not perfect, columns description is not perfect, submission process is not perfect, … Kaggle is not perfect, but it's already great! </p>\n<p>Personally, I enjoyed this competition and hope to see more of this kind in the future!</p>",
      "rawMarkdown": "We all struggled to fit into the constraints of the API and overcome the difficulties of this data. However, people have demonstrated that it was possible to achieve amazing results using a small chunk of the data and/or smart feature pre-processing. Maybe all of this using only Kaggle notebooks.\n\nI understand that sometimes frustration can take over, it happened to me in past competitions. But please keep in mind that Kaggle competitions are not perfect, Kaggle notebooks are not perfect, data is not perfect, columns description is not perfect, submission process is not perfect, ... Kaggle is not perfect, but it's already great! \n\nPersonally, I enjoyed this competition and hope to see more of this kind in the future!",
      "votes": null
    },
    {
      "id": "1144559",
      "postDate": "01/08/2021 14:30:32",
      "content": "<p>I actually enjoyed all of this for one simple reason, Confusions are part of every software project, When we are confused we will try to explore more and more. I agree it's frustrating at multiple times but the results we get when the frustration is over by fixing something recharges you back pretty much. </p>\n<p>I do agree that atleast the code's timeout error should be given before 9 hours, no runtime scrambling is needed IMO and more sample dataset should be given when you have 100M rows, i am pretty sure sacrificing 1M rows might be fine for this. last but not the least, kaggle also had to write the test_iter, so it would have been amazing if they officially shared that code snip which can convert any data chunk to that format which the test_api expects. I know this has a lot of risk if something is wrong etc, the updation info might not reach them, but i am sure people will share that info very quickly as well!</p>",
      "rawMarkdown": "I actually enjoyed all of this for one simple reason, Confusions are part of every software project, When we are confused we will try to explore more and more. I agree it's frustrating at multiple times but the results we get when the frustration is over by fixing something recharges you back pretty much. \n\nI do agree that atleast the code's timeout error should be given before 9 hours, no runtime scrambling is needed IMO and more sample dataset should be given when you have 100M rows, i am pretty sure sacrificing 1M rows might be fine for this. last but not the least, kaggle also had to write the test_iter, so it would have been amazing if they officially shared that code snip which can convert any data chunk to that format which the test_api expects. I know this has a lot of risk if something is wrong etc, the updation info might not reach them, but i am sure people will share that info very quickly as well!",
      "votes": null
    },
    {
      "id": "1144567",
      "postDate": "01/08/2021 14:38:48",
      "content": "<p>I assume some of the obfuscation of the API has to do with preventing probing of the test dataset, so not sure Kaggle will neither released the API code or make it particularly transparent in the next competitions. <br>\nI do agree that it would be nice to kill the submission before 9 hours if memory error occur though </p>",
      "rawMarkdown": "I assume some of the obfuscation of the API has to do with preventing probing of the test dataset, so not sure Kaggle will neither released the API code or make it particularly transparent in the next competitions. \nI do agree that it would be nice to kill the submission before 9 hours if memory error occur though",
      "votes": null
    },
    {
      "id": "1144741",
      "postDate": "01/08/2021 16:34:21",
      "content": "<p>Well not the test API, I meant the way to process the (training) data so that we can get it in the test-set format. This shouldn't be an issue with test set probing, right?</p>",
      "rawMarkdown": "Well not the test API, I meant the way to process the (training) data so that we can get it in the test-set format. This shouldn't be an issue with test set probing, right?",
      "votes": null
    },
    {
      "id": "1145030",
      "postDate": "01/08/2021 20:21:16",
      "content": "<p>I think that is the point of the post though. Yes kaggle is a good platform, but small things could be better and he gave some pretty concrete examples of ways that it was painful for him. This is exactly the kind of feedback and discussion kaggle needs to continue improving.</p>\n<p>That being said I think several of the issues mentioned were specific to this competition and the data which more heavily falls on the organizer and not kaggle themselves. </p>",
      "rawMarkdown": "I think that is the point of the post though. Yes kaggle is a good platform, but small things could be better and he gave some pretty concrete examples of ways that it was painful for him. This is exactly the kind of feedback and discussion kaggle needs to continue improving.\n\nThat being said I think several of the issues mentioned were specific to this competition and the data which more heavily falls on the organizer and not kaggle themselves.",
      "votes": null
    },
    {
      "id": "1145068",
      "postDate": "01/08/2021 21:18:00",
      "content": "<p>\"Incompatibilities between the Kaggle platform and this kind of dataset\" is not exactly the title of a post that suggests improvements. This competition, with all the improvements we may or may not agree on, has been a success, with 3400 participants and a lot of notebooks and discussions. </p>\n<p>For a more constructive level of feedback you may refer for the closing of the 16th place solution from MPWARE. </p>",
      "rawMarkdown": "\"Incompatibilities between the Kaggle platform and this kind of dataset\" is not exactly the title of a post that suggests improvements. This competition, with all the improvements we may or may not agree on, has been a success, with 3400 participants and a lot of notebooks and discussions. \n\nFor a more constructive level of feedback you may refer for the closing of the 16th place solution from MPWARE.",
      "votes": null
    },
    {
      "id": "1145114",
      "postDate": "01/08/2021 22:45:16",
      "content": "<p>By the way,<br>\n\"use a specific previous version's output as input\" - made a post about it 2 months ago =P <a href=\"https://www.kaggle.com/product-feedback/198545\" target=\"_blank\">https://www.kaggle.com/product-feedback/198545</a></p>\n<p>You can also diff committed changes with previous versions. You can use \"Quick Save\" button to quickly commit without actually executing cells. Go to your notebook, click on \"Version X of X\" and you can Select Diff.</p>",
      "rawMarkdown": "By the way,\n\"use a specific previous version's output as input\" - made a post about it 2 months ago =P https://www.kaggle.com/product-feedback/198545\n\nYou can also diff committed changes with previous versions. You can use \"Quick Save\" button to quickly commit without actually executing cells. Go to your notebook, click on \"Version X of X\" and you can Select Diff.",
      "votes": null
    },
    {
      "id": "1145178",
      "postDate": "01/09/2021 00:39:08",
      "content": "<p>I dont understand how one is constructive and the other is not. MPWARE's post is just significantly less verbose and asks for something that has already been requested many times. </p>",
      "rawMarkdown": "I dont understand how one is constructive and the other is not. MPWARE's post is just significantly less verbose and asks for something that has already been requested many times.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1143865,
      "author_name": "tomooinubushi",
      "author_url": "",
      "post_date": "01/08/2021 05:34:20",
      "content": "<p>I have read your topic with a full of sympathy (as many participants do).<br>\nI have been struggling with a lot of submission timeout, memory shortage, and other unknown errors.<br>\nKaggle should upgrade their system with more computational power and more user-friendly way.<br>\nHowever, from the positive side, I could learn how to pre-process a bit dirty time-series data, how to speed up processing,  how to handle data memory efficiently etc. This competition gives me a chance to learn a lot of things, which experienced engineers like you have already acquired.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1144550,
      "author_name": "stecasasso",
      "author_url": "",
      "post_date": "01/08/2021 14:25:19",
      "content": "<p>We all struggled to fit into the constraints of the API and overcome the difficulties of this data. However, people have demonstrated that it was possible to achieve amazing results using a small chunk of the data and/or smart feature pre-processing. Maybe all of this using only Kaggle notebooks.</p>\n<p>I understand that sometimes frustration can take over, it happened to me in past competitions. But please keep in mind that Kaggle competitions are not perfect, Kaggle notebooks are not perfect, data is not perfect, columns description is not perfect, submission process is not perfect, … Kaggle is not perfect, but it's already great! </p>\n<p>Personally, I enjoyed this competition and hope to see more of this kind in the future!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1145030,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "01/08/2021 20:21:16",
          "content": "<p>I think that is the point of the post though. Yes kaggle is a good platform, but small things could be better and he gave some pretty concrete examples of ways that it was painful for him. This is exactly the kind of feedback and discussion kaggle needs to continue improving.</p>\n<p>That being said I think several of the issues mentioned were specific to this competition and the data which more heavily falls on the organizer and not kaggle themselves. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1145068,
          "author_name": "stecasasso",
          "author_url": "",
          "post_date": "01/08/2021 21:18:00",
          "content": "<p>\"Incompatibilities between the Kaggle platform and this kind of dataset\" is not exactly the title of a post that suggests improvements. This competition, with all the improvements we may or may not agree on, has been a success, with 3400 participants and a lot of notebooks and discussions. </p>\n<p>For a more constructive level of feedback you may refer for the closing of the 16th place solution from MPWARE. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1145178,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "01/09/2021 00:39:08",
          "content": "<p>I dont understand how one is constructive and the other is not. MPWARE's post is just significantly less verbose and asks for something that has already been requested many times. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1144559,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "01/08/2021 14:30:32",
      "content": "<p>I actually enjoyed all of this for one simple reason, Confusions are part of every software project, When we are confused we will try to explore more and more. I agree it's frustrating at multiple times but the results we get when the frustration is over by fixing something recharges you back pretty much. </p>\n<p>I do agree that atleast the code's timeout error should be given before 9 hours, no runtime scrambling is needed IMO and more sample dataset should be given when you have 100M rows, i am pretty sure sacrificing 1M rows might be fine for this. last but not the least, kaggle also had to write the test_iter, so it would have been amazing if they officially shared that code snip which can convert any data chunk to that format which the test_api expects. I know this has a lot of risk if something is wrong etc, the updation info might not reach them, but i am sure people will share that info very quickly as well!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1144567,
          "author_name": "stecasasso",
          "author_url": "",
          "post_date": "01/08/2021 14:38:48",
          "content": "<p>I assume some of the obfuscation of the API has to do with preventing probing of the test dataset, so not sure Kaggle will neither released the API code or make it particularly transparent in the next competitions. <br>\nI do agree that it would be nice to kill the submission before 9 hours if memory error occur though </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1144741,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "01/08/2021 16:34:21",
          "content": "<p>Well not the test API, I meant the way to process the (training) data so that we can get it in the test-set format. This shouldn't be an issue with test set probing, right?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1145114,
      "author_name": "returnofsputnik",
      "author_url": "",
      "post_date": "01/08/2021 22:45:16",
      "content": "<p>By the way,<br>\n\"use a specific previous version's output as input\" - made a post about it 2 months ago =P <a href=\"https://www.kaggle.com/product-feedback/198545\" target=\"_blank\">https://www.kaggle.com/product-feedback/198545</a></p>\n<p>You can also diff committed changes with previous versions. You can use \"Quick Save\" button to quickly commit without actually executing cells. Go to your notebook, click on \"Version X of X\" and you can Select Diff.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1143621": "Terribly sorry, I [appear to have written a book](https://en.wikiquote.org/wiki/Blaise_Pascal#Pascal_plus_longue).\n\n# Intro/context\n\nThis was my first Kaggle competition - in fact, I joined Kaggle specifically because of it, attracted by the promise of an interesting educational dataset.\n\nOverall, I am actually very impressed, and often pleasantly surprised, by Kaggle as a platform and a community. It feels like a very nice place to be.\n\nBut on the other hand, this will probably also be my last *code* competition on Kaggle for a good long while, because of the amount of toil (as distinct from work) and confusion that it involved. I estimate that around **80-90%** of the time I spent on this competition was spent on things that decidedly are **not** data science, machine learning, data analysis, or anything that drew my interest initially. Most of that time was spent wrestling with the structure and size of the data, and the many ... *surprising* ways that they interacted with the Kaggle platform. It seems to me (or at least I hope) that these interactions were not at all obvious to the designers of the challenge and/or platform beforehand, so that's why I feel compelled to point them out.\n\nI have seen people on Kaggle use the phrase \"Data Scientists are not Software Engineers\" to try and moderate expectations of what a median Kaggler ought to be able to do in terms of programming, software, knowing what the computer is doing \"under the hood\", etc. This, I think, is very reasonable - the more that you are forced to focus on the implementation details, the less time and mental energy you have to spend on actually doing and/or learning data science.\n\nThat said, I **am** a (former) software engineer. With 10 years of professional experience. 6 of them at Google. If software engineering required this kind of toil, I would have quit a *lot* earlier. I say this not to brag or complain (necessarily). And also not to imply that Kaggle should just do things \"the Google Way\" (in fact, many of the things I love about Kaggle so far would be probably impossible if it was operating as a Google Product). I just want to hopefully give appropriate context to how much of a problem these problems are.\n\n# Analyzing Data\n\nThe main part of the Riiid dataset (the training data itself) is contained a singe .csv file that is around 5GB in size. Some aspects of the data also have a somewhat tricky structure, including:\n\n- a \"content_id\" field that does not, in fact, uniquely identify content. (instead, it's a combination of two unrelated IDs **which intersect in values**). This means that operations like identifying a piece of content, or filtering for types of content, are non-trivial, and (even when implemented correctly *and* efficiently) can frequently consume more processor time and memory than an actual id.\n- two fields (prior_question_had_explanation, prior_question_elapsed_time) whose values relate to **some other** row in the data. Or, actually **some set of other rows** in some cases. Because \"prior_question\" is actually somewhat a misnomer, since these values are associated with \"bundles\" of questions which may or may not contain exactly one question.\n- This is kind of, but not really, time series data - it's really many different independent (or forced-independent) time series, one for each user of the app, forced together by the input file and output API.\n\nWhether these properties of the dataset are, by themselves, a bad idea is out of the scope of this particular post. Instead, like I said, my focus here is how they *interact* with Kaggle as a platform.\n\nThe format of this data challenge - and, from what I've seen, Kaggle in general - leans heavily on using Pandas to process data, and storing state in .csv files:\n\n- the tutorial/API overview focused exclusively on loading the data with Pandas\n- the API itself serves the chunks of test data in Pandas dataframes, and there are no other options offered.\n\nI imagine that normally, this has the purpose of lowering the barrier to entry by limiting the number of tools that a person has to know to get started. But in this case, I think it had the complete opposite effect - the consensus seems to be that it's impossible, or at the very least unwise, to try and deal with this dataset using Pandas alone. So instead, everybody had to go and figure out, individually, some set of tools that actually works for their purpose. Because the starter code and explanation provided for the competition did not adequately start people on the path to dealing with this data successfully.\n\n### Memory\nThe main pain point which arises from the interaction between this dataset, Pandas, and the Kaggle platform is this:\n\n- Pandas tends to be very memory-hungry, often apparently making multiple intermediate copies of the data it's working with.\n    - it's not always easy to predict when and why it will do this\n- Kaggle notebooks are limited to 16GB of RAM, which is roughly 3x the size of the training data\n- When a notebook runs out of those 16GB, **it stops working or restarts, losing intermediate state**\n- What's even worse, when a notebook runs out of memory in non-interactive mode (i.e. on submit), **it hangs until it runs out of time** instead of stopping and reporting the problem.\n    - Apparently, this has been a \"[known issue, will not fix](https://www.kaggle.com/product-feedback/71176)\" for at least two years. The advice is basically to test it in interactive mode first. But this isn't very practical if, say, the memory failure only happens 5 hours into the processing - which is quite likely and (for me) common when trying to deal with this dataset.\n        - Surely, if it's possible to *know* I'm out of memory in interactive mode, it should be possible to at least *guess* this is happening in the other mode, and alert me to this fact, so that I don't have to interactively watch my batch submission to manually guess whether it has actually failed 3 hours ago?\n    - This issue alone - which is undocumented, except for that two year old forum post - cost me a couple of days **just trying to understand what is happening**\n    - Thrashing memory for 9 hours because it ran out of memory in the  first 10 minutes can't possibly be a good use of resources?..\n\n    \nOf course, this problem isn't limited to Pandas. Fundamentally, 5GB of data is sitting in a single file, and it needs to go into memory, and get manipulated, all without ever going over 16GB total in memory, **and** without going over 9 hours of processing time for whatever operation you're trying to do.\n\nThis problem is exacerbated when complex operations are, in fact, necessary to get information out of the data which is present but misaligned, as with the prior_question fields. In the end I did manage to associate these fields with the appropriate row in the table, but I had to find a very precise way of shuffling the data between Pandas and Datatable to avoid their respective weak points.\n\nIf these rows had not been shifted in the first place - if the training data had this_question_had_explanation and this_question_elapsed_time - I think many more people would have been much less confused, and more able to come up with interesting ways of using that data. Just search for \"prior_question\" in the discussion forums and marvel at the length and confusion of the posts.\n\nI can only imagine that the motivation behind shifting the data was to make it more clear that these particular fields couldn't be used directly in training, but again, I think it created way more confusion than it resolved. I think any number of alternative solutions would have worked so much better:\n\n- Not shifting, and just treating those fields the same exact way as answered_correctly in the test data  (and providing a clear example of using this test data)\n- shifting, but providing an additional field(or fields) which tracks what the prior_question **was**\n- shifting, but having a less complicated and better-explained scheme about what \"prior\" means\n    - Actually have people on hand to explain this in the forums, instead of leaving us to guess and experiment\n\n### Notebooks\n\nThis is a much smaller issue than the memory/data size/data format interaction. But I often found the options for managing, versioning and chaining notebooks to be quite limiting.\n\nI ended up making 40+ notebooks for this data challenge. Some are just exploratory, some are chained together in that one notebook produces output which another notebook then uses. I think that if the data was more manageable, the number and variety of notebooks would probably also go down and be more manageable. But as is, I needed to keep up with many different independent bits, all of which had several different (potentially broken) versions.\n\nSpecific things I found lacking:\n\n- If I have to find a particular notebook, my only option is to stare at a flat list of names and hope I remember what I named it. This is exacerbated by the quite-short character limit on notebook names.\n- There is clearly a version control system in place, but frustratingly few capabilities are exposed to us. In particular, I found myself wanting (but unable) to:\n    - revert to a previous version, or even just discard uncommitted changes\n        - or heck, even just diff my uncommitted changes with the committed version\n    - use a specific previous version's output as input (e.g. either because the current one is broken - and will be for another 5 hours as the fix re-runs)\n    - fork from a previous version (or even just from the committed current version I'm looking at when I press \"fork\", instead of the draft)\n- I also would very much appreciate some kind of graph view of my related notebooks: which ones are forked from which? which ones feed data into which? I honestly couldn't tell you right now what my 40+ notebooks do and how that happened. but if I had a graph, I might.\n\n# Submission process\n\nIt took me 12 submissions over 3 days just to get by baseline code to submit, and after that it took me about 10 more to add just one more bit of logic to it. After that, I ran out of time, so I really didn't get to do a lot of interesting stuff. Then again, running out of time might have better for my sanity than the alternative.\n\n### Debugging and API shape\n\nI think lots of people have already talked about this extensively, so I'll try to keep at least this section brief:\n\n- The fact that there is no information about why a submission failed is absolutely unworkable\n    - It's not even easy to tell **how fast** my submission failed unless I am watching the submission run (and then also recording it myself somewhere) - e.g. did this submission from a day ago fail after 5 minutes because I have a bug, or after 9 hours because it timed out? \n- This could have been mitigated by a robust test set which includes all known edge cases. Instead, the set of data that's visible when we are testing the API seems to have exactly **one** edge case (new user), and is conspicuously missing another edge case which the organizers already knew was important (lectures) - they put the (one-sentence) description of this edge case in bold in their starter notebook.\n    - Possibly even a small synthetic data set which covers lots of edge cases would have been a major improvement\n        - of course, these edge cases would have to be found first before the competition starts. Perhaps through internal testing?..\n- This is all further  exacerbated by the fact that the API shape has several gotchas.\n    - The main one is that the predictions and the subsequent \"correct answers to your previous submission\" have a different shape: one has -1 for lectures, the other one MUST skip lectures. This was not explicitly explained by organizers, and again, not testable.\n    - The part where the correct answers are a serialized bit of python shoved into some arbitrary single field in an unrelated dataframe... Actually was far less problematic than I thought it would be. But it *is* a strong indicator that this API was not designed for this challenge - that somebody took a well-thought-out API for a different kind of problem (actual time series data, without a need to track state of previous submissions) and tried to adapt it by making the smallest possible number of changes. Instead of designing an API to specifically fit this dataset.\n        - I can see how this decision, too, may have been made in the name of simplicity (people are already used to this API...), but it only made things more confusing and complicated.\n- In these circumstances, the 5 submission per day limit is extremely frustrating. For that matter, why do failed submissions **ever** count toward that limit?\n\n### State\n\nOhhh boy, and I thought the prior_question fields were problematic when I was trying to process training data.\n\nThis challenge is - or ought to be - a natural fit for tracking user state. Because we are trying to predict user actions based on previous user actions (well, and the app's choices based on user actions, but that's a topic for a whole different post). However, **previous user state is split up in such a way that it's extremely hard to put it back together**.\n\nWe have some data (answered_correctly, content_id, ...) associated with **the previous chunk of test data** (or, **just the first time**, from the state at the end of training). And some data (prior_question fields) associated with **that user's previous most recent answer**, which may come from *some* other chunk of test data, or the training data, **at any point in the test process**.\n\nSo when I'm trying to actually collate all this and understand, for example, **whether a user answered their previous question correctly, how long it took them to answer that question, what that question was, and whether they received feedback**, how do I do that?.. Suppose I am tracking general \"last known user state\" and updating it as I get information from the various sources.\n\n- **At what point does the prior_question data actually line up with the answered_correctly data, so they both refer to the same question?**\n- Suppose I also care about question counts, and dividing various totals by the number of questions the user has seen. Can I guarantee that just one question count is actually up-to-date for all the fields simultaneously?..\n\nThis might actually make for an interesting Software Engineer interview question. But in simplified form, on a whiteboard, and without the complete lack of feedback. \n\nI know for a fact that I have at least one bug in my logic for dealing with this. But whenever I touched that code, my entire submission would start failing again. And I just did not have enough submission budget to experiment with it.\n\n### Time\n\nWe are given 9 hours minus 15 minutes (according to the instructions, the API takes up 15 minutes for itself) to predict 2.5 million data points. That means about 80-100 predictions a second (depending on load/spin-up time of my own data). This is motivated by saying that these predictions should be usable in a real-time setting. But no real-time setting needs predictions to be made in 0.01 of a second, while keeping the context of the entire userbase in RAM. Consider that a reasonable screen refresh rate is 30 frames a second. Making predictions at 3x that speed is pretty useless optimization. The reality, of course, is that 9 hours is the standard Kaggle notebook running time, and it couldn't be changed for this competition. In that case, though, the size of the test dataset should have been changed.\n\nOf course, making 100 predicions a second is only a problem when you want to keep complex user state and update it with elaborate logic. \n\nI guess you could argue that severely limiting the feasibility of tracking user state actually \"simplifies\" the challenge somewhat by encouraging *everyone* to not track complex state manually; instead, either rely on your model to implicitly track an approximation of it, or just throw out annoying state altogether. But this kind of \"simplification\" is sort of like simplifying [the process of searching for your keys in the dark](https://en.wikipedia.org/wiki/Streetlight_effect) by smashing some of the streetlights, in order to make the feasible search space smaller.\n\n# Conclusions\n\nThe ideal (given unlimited resources) solution to situations like this, I think, would be to:\n\n- Store data like this in a database at the outset, not a text file\n- Provide examples of using appropriate tools for the data format and optimization goal (i.e. for predicting correctness in small chunks - probably not (just) pandas)\n- Don't structure the data in a way that requires either (1) esoteric processing or (2) effectively ignoring most of the usefulness of that part of the data\n- Design the (size and structure) of the evaluation part of the competition to be reasonably compatible with Kaggle's performance constraints\n- Design the submission API to fit the specific dataset\n- Provide robust tools and/or data sets for diagnosing and debugging problems \n    - Maybe something like the [emulator](https://www.kaggle.com/its7171/time-series-api-iter-test-emulator) made by one of the community members, but actually guaranteed to match the API\n- [Dogfood](https://en.wikipedia.org/wiki/Eating_your_own_dog_food) and make adjustments when it's still easy/possible to change the format of the challenge\n    - And I don't mean \"do a run the pre-existing model around which this competition was designed\". I mean get other people, who were not involved in designing the challenge, to earnestly try to do some interesting stuff with the data and go through the submission flow on their own.",
    "1143865": "I have read your topic with a full of sympathy (as many participants do).\nI have been struggling with a lot of submission timeout, memory shortage, and other unknown errors.\nKaggle should upgrade their system with more computational power and more user-friendly way.\nHowever, from the positive side, I could learn how to pre-process a bit dirty time-series data, how to speed up processing,  how to handle data memory efficiently etc. This competition gives me a chance to learn a lot of things, which experienced engineers like you have already acquired.",
    "1144550": "We all struggled to fit into the constraints of the API and overcome the difficulties of this data. However, people have demonstrated that it was possible to achieve amazing results using a small chunk of the data and/or smart feature pre-processing. Maybe all of this using only Kaggle notebooks.\n\nI understand that sometimes frustration can take over, it happened to me in past competitions. But please keep in mind that Kaggle competitions are not perfect, Kaggle notebooks are not perfect, data is not perfect, columns description is not perfect, submission process is not perfect, ... Kaggle is not perfect, but it's already great! \n\nPersonally, I enjoyed this competition and hope to see more of this kind in the future!",
    "1144559": "I actually enjoyed all of this for one simple reason, Confusions are part of every software project, When we are confused we will try to explore more and more. I agree it's frustrating at multiple times but the results we get when the frustration is over by fixing something recharges you back pretty much. \n\nI do agree that atleast the code's timeout error should be given before 9 hours, no runtime scrambling is needed IMO and more sample dataset should be given when you have 100M rows, i am pretty sure sacrificing 1M rows might be fine for this. last but not the least, kaggle also had to write the test_iter, so it would have been amazing if they officially shared that code snip which can convert any data chunk to that format which the test_api expects. I know this has a lot of risk if something is wrong etc, the updation info might not reach them, but i am sure people will share that info very quickly as well!",
    "1144567": "I assume some of the obfuscation of the API has to do with preventing probing of the test dataset, so not sure Kaggle will neither released the API code or make it particularly transparent in the next competitions. \nI do agree that it would be nice to kill the submission before 9 hours if memory error occur though",
    "1144741": "Well not the test API, I meant the way to process the (training) data so that we can get it in the test-set format. This shouldn't be an issue with test set probing, right?",
    "1145030": "I think that is the point of the post though. Yes kaggle is a good platform, but small things could be better and he gave some pretty concrete examples of ways that it was painful for him. This is exactly the kind of feedback and discussion kaggle needs to continue improving.\n\nThat being said I think several of the issues mentioned were specific to this competition and the data which more heavily falls on the organizer and not kaggle themselves.",
    "1145068": "\"Incompatibilities between the Kaggle platform and this kind of dataset\" is not exactly the title of a post that suggests improvements. This competition, with all the improvements we may or may not agree on, has been a success, with 3400 participants and a lot of notebooks and discussions. \n\nFor a more constructive level of feedback you may refer for the closing of the 16th place solution from MPWARE.",
    "1145114": "By the way,\n\"use a specific previous version's output as input\" - made a post about it 2 months ago =P https://www.kaggle.com/product-feedback/198545\n\nYou can also diff committed changes with previous versions. You can use \"Quick Save\" button to quickly commit without actually executing cells. Go to your notebook, click on \"Version X of X\" and you can Select Diff.",
    "1145178": "I dont understand how one is constructive and the other is not. MPWARE's post is just significantly less verbose and asks for something that has already been requested many times."
  },
  "source": "meta"
}