{
  "id": 62883,
  "title": "Parallel loop in Python",
  "url": "/competitions/trackml-particle-identification/discussion/62883",
  "author_name": "",
  "post_date": "2018-08-08T13:05:46.252041500Z",
  "votes": 14,
  "comment_count": 70,
  "views": 0,
  "content": "<p>As many noticed, it is possible to speed up submission computation using parallel processing, with a master process and a number of worker processes.  Python supports it, but there are few tricks to make it efficient.</p>\n\n<ol>\n<li><p>Do not pass data to the workers, rather have each worker load its data from disk.  Indeed, passing data means the master process must load it first, and this can become a real bottleneck for the whole computation.</p></li>\n<li><p>Make sure each worker works on a single event.  Indeed, Python Pool assigns more than one task to each worker with its default settings. This can result in some of your processors becoming idle while events remain to be processed.</p></li>\n</ol>\n\n<p>Code to do it in Python 3.6 is given below.  It assumes you implement the computation of one event submission  as a single function <code>work_sub</code> that gets one parameter, namely the event number.  You can pass additional parameters but be careful, all data you pass there can create a bottleneck for the whole computation, stick to simple parameters like numerical values.  Do not pass the event data.</p>\n\n<p>Then in pseudo code:</p>\n\n<pre><code>from multiprocessing import Pool\n\ndef work_sub(param):\n    (i, &lt;other parameters&gt;) = param\n    &lt;load data for event i, compute its submission, then save to disk&gt;\n    return i\n\nparams = [(i, &lt;other parameters&gt;) for i in range(125)]\n\nn_proc = &lt;number of physical cores on your machine&gt;\n\nuse_parallel = 1\nif use_parallel: \n    pool = Pool(processes=n_proc, maxtasksperchild=1)\n    ls   = pool.map( work_sub, params, chunksize=1 )\n    pool.close()\nelse:\n    ls = [work_sub(param) for param in params]\n</code></pre>\n\n<p>Once completed, you just have to load each event submission and concatenate them to get your final submission.  Make sure you have the right event_id in each of the event submission.</p>\n\n<p>The <code>use_parallel</code>  is for you to check your code before running it for real: set <code>use_parallel</code> to False, make sure the code runs fine for one event, then restart with <code>use_parallel</code>  set to True.</p>\n\n<p>I recommend you print some progress in the <code>work_sub</code> function, or even use <code>tqdm</code> to monitor how your computation goes.  </p>",
  "messages": [
    {
      "id": "367749",
      "postDate": "08/08/2018 13:05:46",
      "content": "<p>As many noticed, it is possible to speed up submission computation using parallel processing, with a master process and a number of worker processes.  Python supports it, but there are few tricks to make it efficient.</p>\n\n<ol>\n<li><p>Do not pass data to the workers, rather have each worker load its data from disk.  Indeed, passing data means the master process must load it first, and this can become a real bottleneck for the whole computation.</p></li>\n<li><p>Make sure each worker works on a single event.  Indeed, Python Pool assigns more than one task to each worker with its default settings. This can result in some of your processors becoming idle while events remain to be processed.</p></li>\n</ol>\n\n<p>Code to do it in Python 3.6 is given below.  It assumes you implement the computation of one event submission  as a single function <code>work_sub</code> that gets one parameter, namely the event number.  You can pass additional parameters but be careful, all data you pass there can create a bottleneck for the whole computation, stick to simple parameters like numerical values.  Do not pass the event data.</p>\n\n<p>Then in pseudo code:</p>\n\n<pre><code>from multiprocessing import Pool\n\ndef work_sub(param):\n    (i, &lt;other parameters&gt;) = param\n    &lt;load data for event i, compute its submission, then save to disk&gt;\n    return i\n\nparams = [(i, &lt;other parameters&gt;) for i in range(125)]\n\nn_proc = &lt;number of physical cores on your machine&gt;\n\nuse_parallel = 1\nif use_parallel: \n    pool = Pool(processes=n_proc, maxtasksperchild=1)\n    ls   = pool.map( work_sub, params, chunksize=1 )\n    pool.close()\nelse:\n    ls = [work_sub(param) for param in params]\n</code></pre>\n\n<p>Once completed, you just have to load each event submission and concatenate them to get your final submission.  Make sure you have the right event_id in each of the event submission.</p>\n\n<p>The <code>use_parallel</code>  is for you to check your code before running it for real: set <code>use_parallel</code> to False, make sure the code runs fine for one event, then restart with <code>use_parallel</code>  set to True.</p>\n\n<p>I recommend you print some progress in the <code>work_sub</code> function, or even use <code>tqdm</code> to monitor how your computation goes.  </p>",
      "rawMarkdown": "As many noticed, it is possible to speed up submission computation using parallel processing, with a master process and a number of worker processes.  Python supports it, but there are few tricks to make it efficient.\n\n1. Do not pass data to the workers, rather have each worker load its data from disk.  Indeed, passing data means the master process must load it first, and this can become a real bottleneck for the whole computation.\n\n2. Make sure each worker works on a single event.  Indeed, Python Pool assigns more than one task to each worker with its default settings. This can result in some of your processors becoming idle while events remain to be processed.\n\nCode to do it in Python 3.6 is given below.  It assumes you implement the computation of one event submission  as a single function `work_sub` that gets one parameter, namely the event number.  You can pass additional parameters but be careful, all data you pass there can create a bottleneck for the whole computation, stick to simple parameters like numerical values.  Do not pass the event data.\n\nThen in pseudo code:\n\n    from multiprocessing import Pool\n    \n    def work_sub(param):\n        (i,",
      "votes": null
    },
    {
      "id": "367780",
      "postDate": "08/08/2018 14:08:15",
      "content": "<blockquote>\n  <p>passing data means the master process must load it first, and this can become a real bottleneck for the whole computation.</p>\n</blockquote>\n\n<p>Do you know a time for passing data? Suppose 1 dbscan is about 1 second. Is it comparable to it? I pass 'hits' values and other params. I thought passing is fast. :(</p>",
      "rawMarkdown": "&gt; passing data means the master process must load it first, and this can become a real bottleneck for the whole computation.\n\nDo you know a time for passing data? Suppose 1 dbscan is about 1 second. Is it comparable to it? I pass 'hits' values and other params. I thought passing is fast. :(",
      "votes": null
    },
    {
      "id": "367783",
      "postDate": "08/08/2018 14:22:49",
      "content": "<p>The problem is not that passing is fast or not.  The problem is that python cannot execute more than one thread at a time because of the GIL.  It means that the master can pass data to only one worker at a time, making it non parallel...  The effect is worse when your worker runs fast, or if you have many workers.</p>",
      "rawMarkdown": "The problem is not that passing is fast or not.  The problem is that python cannot execute more than one thread at a time because of the GIL.  It means that the master can pass data to only one worker at a time, making it non parallel...  The effect is worse when your worker runs fast, or if you have many workers.",
      "votes": null
    },
    {
      "id": "367790",
      "postDate": "08/08/2018 14:42:54",
      "content": "<p>The worst I can imagine is to load all events data in the master, then pass it to each worker.  Then the worker code selects the relevant part of it.  </p>\n\n<p>With this pattern you load the 125 events data <code>n_proc + 1</code> times, overloading your machine memory, and slowing it dramatically.</p>",
      "rawMarkdown": "The worst I can imagine is to load all events data in the master, then pass it to each worker.  Then the worker code selects the relevant part of it.  \n\nWith this pattern you load the 125 events data `n_proc + 1` times, overloading your machine memory, and slowing it dramatically.",
      "votes": null
    },
    {
      "id": "367800",
      "postDate": "08/08/2018 15:15:03",
      "content": "<p>\"You can pass additional parameters but be careful, all data you pass there can create a bottleneck for the whole computation, stick to simple parameters like numerical values. \" - Ok, but then if you are required to get the data for event i, you are still required to access the event data by doing something like <code>df.loc[df['event']==i]</code>. So you are forced to pass in <code>df</code>, unless you make <code>df</code> a global variable</p>",
      "rawMarkdown": "\"You can pass additional parameters but be careful, all data you pass there can create a bottleneck for the whole computation, stick to simple parameters like numerical values. \" - Ok, but then if you are required to get the data for event i, you are still required to access the event data by doing something like `df.loc[df['event']==i]`. So you are forced to pass in `df`, unless you make `df` a global variable",
      "votes": null
    },
    {
      "id": "367811",
      "postDate": "08/08/2018 15:36:31",
      "content": "<p>You are describing one of the worst, if not the worst, way of implementing parallelism: you make all event data copied in each worker.</p>\n\n<blockquote>\n  <p>you are still required to access the event data by doing something like df.loc[df['event']==i]</p>\n</blockquote>\n\n<p>Absolutely not, you can load each event data independently in each worker.  For instance, here is the start of my worker code:</p>\n\n<pre><code>def get_event(i):\n    prefix = 'event000000'\n    if i &lt; 100:\n        prefix = prefix +'0'\n    if i &lt; 10:\n        prefix = prefix +'0'\n    return prefix + str(i)\n\nbase_path = '/home/jfpuget/Kaggle/TrackML/'\n\ndef work_sub(param):\n    (i, &lt;other parameters&gt;) = param\n\n    event = get_event(i)\n    print('event:', event)\n    hits = pd.read_csv(base_path+'input/test/'+event + '-hits.csv')\n</code></pre>",
      "rawMarkdown": "You are describing one of the worst, if not the worst, way of implementing parallelism: you make all event data copied in each worker.\n\n&gt; you are still required to access the event data by doing something like df.loc[df['event']==i]\n\nAbsolutely not, you can load each event data independently in each worker.  For instance, here is the start of my worker code:\n\n    def get_event(i):\n        prefix = 'event000000'\n        if i &lt; 100:\n            prefix = prefix +'0'\n        if i &lt; 10:\n            prefix = prefix +'0'\n        return prefix + str(i)\n    \n    base_path = '/home/jfpuget/Kaggle/TrackML/'\n    \n    def work_sub(param):\n        (i,",
      "votes": null
    },
    {
      "id": "367818",
      "postDate": "08/08/2018 15:49:24",
      "content": "<p>So you have solved it by reading that data in CSV form, as opposed to making one really big data frame. What if am I restricted to the large dataframe with event ID's in it? I feel that I become stuck because I must somehow only get <code>df.loc[df['event']==i]</code> yet I can't pass <code>df</code>! Hm. Would this make do as a workaround, or would it still make bottlenecks?</p>\n\n<pre><code>def get_hits(i):\n    df_temp = df.loc[df['event']==i]\n    return df_temp\n\n\ndef work_sub(param):\n    (i, &lt;other parameter&gt;) = param\n    hits = get_hits(i)\n</code></pre>\n\n<p>I predict this will still make a bottleneck, however, because to execute <code>get_hits</code> the worker needs to read in <code>df</code>. Therefore I fear the only way to make this work is if your dataset can be read one-by-one, like as if it were stored in CSV format earlier so you can read it in by the event_id. If it's in one large dataframe, it seems I'm screwed . . !</p>",
      "rawMarkdown": "So you have solved it by reading that data in CSV form, as opposed to making one really big data frame. What if am I restricted to the large dataframe with event ID's in it? I feel that I become stuck because I must somehow only get `df.loc[df['event']==i]` yet I can't pass `df`! Hm. Would this make do as a workaround, or would it still make bottlenecks?\n\n    def get_hits(i):\n        df_temp = df.loc[df['event']==i]\n        return df_temp\n\n\n    def work_sub(param):\n        (i,",
      "votes": null
    },
    {
      "id": "367821",
      "postDate": "08/08/2018 15:52:00",
      "content": "<p>I've taken measurements in my case. The overhead of parallelization is about 8% (including parameter passing). A lot, but not deadly.</p>\n\n<p>And thanks for this info! I didn't think about this overhead.</p>",
      "rawMarkdown": "I've taken measurements in my case. The overhead of parallelization is about 8% (including parameter passing). A lot, but not deadly.\n\nAnd thanks for this info! I didn't think about this overhead.",
      "votes": null
    },
    {
      "id": "367822",
      "postDate": "08/08/2018 15:52:06",
      "content": "<p>Do you get that your way means the big dataframe is copied in each worker?  If you don't then I suggest you study how parallelism works in Python.  If you do then I think you get why the way I propose is better.</p>\n\n<p>Edit: why do you want to load all event data, concatenate it into a gigantic data frame, then have it split back into each worker?  </p>",
      "rawMarkdown": "Do you get that your way means the big dataframe is copied in each worker?  If you don't then I suggest you study how parallelism works in Python.  If you do then I think you get why the way I propose is better.\n\nEdit: why do you want to load all event data, concatenate it into a gigantic data frame, then have it split back into each worker?",
      "votes": null
    },
    {
      "id": "367835",
      "postDate": "08/08/2018 16:15:12",
      "content": "<p>I am learning from your example because I'm implementing parallelism at my work, and I read Kaggle posts for learning. I'm not given each event data as an individual CSV, and I am trying to use your tips to change my current parallelism implementation to make it run faster. I <em>start</em> with one gigantic data frame.</p>\n\n<p>According to this post, <a href=\"https://stackoverflow.com/questions/33612935/large-pandas-dataframe-parallel-processing\">https://stackoverflow.com/questions/33612935/large-pandas-dataframe-parallel-processing</a> it appears that I am forced to let the parallel package pickle my data frame each time a worker wants to access it.</p>\n\n<p>According to this post, <a href=\"https://stackoverflow.com/questions/40357434/pandas-df-iterrow-parallelization\">https://stackoverflow.com/questions/40357434/pandas-df-iterrow-parallelization</a> a partial workaround is to split it up into the number of processes I want to have.</p>\n\n<p>Also recommended was a package called Dask. </p>\n\n<p>For now however, because my dataframe is small (only about 50,000 rows), I think I might change my code to split the dataframe as recommended by the second post.</p>",
      "rawMarkdown": "I am learning from your example because I'm implementing parallelism at my work, and I read Kaggle posts for learning. I'm not given each event data as an individual CSV, and I am trying to use your tips to change my current parallelism implementation to make it run faster. I *start* with one gigantic data frame.\n\nAccording to this post, https://stackoverflow.com/questions/33612935/large-pandas-dataframe-parallel-processing it appears that I am forced to let the parallel package pickle my data frame each time a worker wants to access it.\n\nAccording to this post, https://stackoverflow.com/questions/40357434/pandas-df-iterrow-parallelization a partial workaround is to split it up into the number of processes I want to have.\n\nAlso recommended was a package called Dask. \n\nFor now however, because my dataframe is small (only about 50,000 rows), I think I might change my code to split the dataframe as recommended by the second post.",
      "votes": null
    },
    {
      "id": "367841",
      "postDate": "08/08/2018 16:30:07",
      "content": "<p>I suggest you first write each event data in a separate file on disk, then you can implement what I describe, loading only the relevant file into each worker.</p>\n\n<p>The first post you point to is also documenting the main issue with joblib: you have to load all data in the master, then it is copied into each worker.  That's why I am not using joblib and rather use Python parallel processing directly.</p>",
      "rawMarkdown": "I suggest you first write each event data in a separate file on disk, then you can implement what I describe, loading only the relevant file into each worker.\n\nThe first post you point to is also documenting the main issue with joblib: you have to load all data in the master, then it is copied into each worker.  That's why I am not using joblib and rather use Python parallel processing directly.",
      "votes": null
    },
    {
      "id": "367843",
      "postDate": "08/08/2018 16:31:29",
      "content": "<p>Parallelism overhead is so small it is not measurable with my code ;)  I agree that 8% is not deadly, I guess you don't use many parallel processes.  </p>",
      "rawMarkdown": "Parallelism overhead is so small it is not measurable with my code ;)  I agree that 8% is not deadly, I guess you don't use many parallel processes.",
      "votes": null
    },
    {
      "id": "367862",
      "postDate": "08/08/2018 17:33:59",
      "content": "<p>This seem a bit more elegant than my solution which was to process all events without saving each to disk on a 72 core aws instance with about 144gb of ram. (c5.18xlarge) Although it was able to process the whole submission in 3 hours which was just over twice the time it took to process one event locally.</p>",
      "rawMarkdown": "This seem a bit more elegant than my solution which was to process all events without saving each to disk on a 72 core aws instance with about 144gb of ram. (c5.18xlarge) Although it was able to process the whole submission in 3 hours which was just over twice the time it took to process one event locally.",
      "votes": null
    },
    {
      "id": "367873",
      "postDate": "08/08/2018 18:12:00",
      "content": "<p>Note that what I suggest makes sense for the specific computation we have here, given each event data is in separate files.  If you are discussing parallel processing of data frames in general, then it is different.  </p>",
      "rawMarkdown": "Note that what I suggest makes sense for the specific computation we have here, given each event data is in separate files.  If you are discussing parallel processing of data frames in general, then it is different.",
      "votes": null
    },
    {
      "id": "367874",
      "postDate": "08/08/2018 18:13:50",
      "content": "<p>&gt; I guess you don't use many parallel processes. </p>\n\n<p>You're right. I use 3 processes (sometimes on different machines). If use all 4 cores, all becomes slow and I can't do anything else.</p>",
      "rawMarkdown": "&gt; I guess you don't use many parallel processes. \n\nYou're right. I use 3 processes (sometimes on different machines). If use all 4 cores, all becomes slow and I can't do anything else.",
      "votes": null
    },
    {
      "id": "367876",
      "postDate": "08/08/2018 18:15:44",
      "content": "<p>See, your way is not enabling you to use your 4 cores...  I bet you would be able to use them if you move to the way I describe above.</p>\n\n<p>Your overhead is way higher than you think, you should measure it when running with 4 workers (one per core).</p>",
      "rawMarkdown": "See, your way is not enabling you to use your 4 cores...  I bet you would be able to use them if you move to the way I describe above.\n\nYour overhead is way higher than you think, you should measure it when running with 4 workers (one per core).",
      "votes": null
    },
    {
      "id": "367884",
      "postDate": "08/08/2018 18:46:24",
      "content": "<p>thank you for this post! It looks easier then <a href=\"https://pythonhosted.org/joblib/parallel.html\">https://pythonhosted.org/joblib/parallel.html</a> I tried to implement before... we'll see</p>",
      "rawMarkdown": "thank you for this post! It looks easier then https://pythonhosted.org/joblib/parallel.html I tried to implement before... we'll see",
      "votes": null
    },
    {
      "id": "367927",
      "postDate": "08/08/2018 21:03:39",
      "content": "<p>I was confused by the parameter 'maxtasksperchild=1'. The documentation says ' A frequent pattern found in other systems (such as Apache, mod_wsgi, etc) to free resources held by workers is to allow a worker within a pool to complete only a set amount of work before being exiting, being cleaned up and a new process spawned to replace the old one.'\nI think it can take to spawn a new process. Why not to use an existing process? Memory fragmetation etc?</p>",
      "rawMarkdown": "I was confused by the parameter 'maxtasksperchild=1'. The documentation says ' A frequent pattern found in other systems (such as Apache, mod_wsgi, etc) to free resources held by workers is to allow a worker within a pool to complete only a set amount of work before being exiting, being cleaned up and a new process spawned to replace the old one.'\nI think it can take to spawn a new process. Why not to use an existing process? Memory fragmetation etc?",
      "votes": null
    },
    {
      "id": "367936",
      "postDate": "08/08/2018 21:31:05",
      "content": "<p>This is what I use for all the version of Clusterer() out there. I doesn't seem to use much memory. I guess if you had more cores you would have more memory too?</p>\n\n<pre><code>def worker(event_id, hits, cells):\n    model = Clusterer()\n    labels = model.predict(hits)\n    print('Event ID: ', event_id)\n    # Prepare submission for an event\n    one_submission = create_one_event_submission(event_id, hits, labels)\n    return one_submission\n\nfrom multiprocessing import Pool\n\npool = Pool(processes=6)\ntest_dataset_submissions = pool.starmap(worker,load_dataset(path_to_test, parts=['hits', 'cells']))\n\n# Create submission file\nsubmission = pd.concat(test_dataset_submissions, axis=0)\nsubmission.to_csv('submission.csv', index=False)\n</code></pre>",
      "rawMarkdown": "This is what I use for all the version of Clusterer() out there. I doesn't seem to use much memory. I guess if you had more cores you would have more memory too?\n\n\n    def worker(event_id, hits, cells):\n        model = Clusterer()\n        labels = model.predict(hits)\n        print('Event ID: ', event_id)\n        # Prepare submission for an event\n        one_submission = create_one_event_submission(event_id, hits, labels)\n        return one_submission\n    \n    from multiprocessing import Pool\n    \n    pool = Pool(processes=6)\n    test_dataset_submissions = pool.starmap(worker,load_dataset(path_to_test, parts=['hits', 'cells']))\n    \n    # Create submission file\n    submission = pd.concat(test_dataset_submissions, axis=0)\n    submission.to_csv('submission.csv', index=False)",
      "votes": null
    },
    {
      "id": "367951",
      "postDate": "08/08/2018 23:26:00",
      "content": "<p>I am new to python, sorry for this stupid questions: \n1. I guess @CPMP suggests to place loader in the worker, it speeds up, and he is using his hits = pd.read_csv() instead of trackml libruary </p>\n\n<ol>\n<li>pool.starmap(worker,load_dataset(path_to_test, parts=['hits', 'cells'])) \nshould it be worker(event_id, hits, cells = load_dataset(path_to_test, parts=['hits', 'cells'])) </li>\n</ol>\n\n<p>should it be the for loop for events from 1 to 125 somewhere? </p>",
      "rawMarkdown": "I am new to python, sorry for this stupid questions: \n1. I guess @CPMP suggests to place loader in the worker, it speeds up, and he is using his hits = pd.read_csv() instead of trackml libruary \n\n2. pool.starmap(worker,load_dataset(path_to_test, parts=['hits', 'cells'])) \nshould it be worker(event_id, hits, cells = load_dataset(path_to_test, parts=['hits', 'cells'])) \n\nshould it be the for loop for events from 1 to 125 somewhere?",
      "votes": null
    },
    {
      "id": "367960",
      "postDate": "08/09/2018 00:27:28",
      "content": "<p>The load time isn't my bottleneck, so I'm loading from the main part, and if all the process start loading files at once it could create another slowdown... you can never win with parallelism.\nload_dataset(path_to_test, parts=['hits', 'cells']) gives you a list of tuples like this:</p>\n\n<pre><code>[(0,hits,cells),(1,hits,cells),...]\n</code></pre>\n\n<p>so if you just pass the tuple to the worker function you would have to handle that some way. Or you pass with starmap and it sends 3 things to the worker instead.\nHope that clears it up.</p>",
      "rawMarkdown": "The load time isn't my bottleneck, so I'm loading from the main part, and if all the process start loading files at once it could create another slowdown... you can never win with parallelism.\nload_dataset(path_to_test, parts=['hits', 'cells']) gives you a list of tuples like this:\n\n    [(0,hits,cells),(1,hits,cells),...]\n\nso if you just pass the tuple to the worker function you would have to handle that some way. Or you pass with starmap and it sends 3 things to the worker instead.\nHope that clears it up.",
      "votes": null
    },
    {
      "id": "368056",
      "postDate": "08/09/2018 06:43:58",
      "content": "<p>This parameter forces to stop each worker process after it has processed one event, thereby freeing all resources associated with it.  If you don't use it, then you keep <code>n_proc</code> workers resources until the end of computation.  It does not make much difference till the number of remaining events is less than <code>n_proc</code>. At that point the number of workers starts decreasing with this setting, rather than having idle workers.</p>",
      "rawMarkdown": "This parameter forces to stop each worker process after it has processed one event, thereby freeing all resources associated with it.  If you don't use it, then you keep `n_proc` workers resources until the end of computation.  It does not make much difference till the number of remaining events is less than `n_proc`. At that point the number of workers starts decreasing with this setting, rather than having idle workers.",
      "votes": null
    },
    {
      "id": "368057",
      "postDate": "08/09/2018 06:50:50",
      "content": "<p>&gt; The load time isn't my bottleneck</p>\n\n<p>How do you know that?</p>\n\n<p>Anyway, your way of doing thing is the one that is slowest as I explained, because you force all data to be loaded by the master, then copied to each worker.  It does not matter much if you use a small number of workers, but will degrade computation time if you use many.</p>",
      "rawMarkdown": "&gt; The load time isn't my bottleneck\n\nHow do you know that?\n\nAnyway, your way of doing thing is the one that is slowest as I explained, because you force all data to be loaded by the master, then copied to each worker.  It does not matter much if you use a small number of workers, but will degrade computation time if you use many.",
      "votes": null
    },
    {
      "id": "368163",
      "postDate": "08/09/2018 11:46:27",
      "content": "<p>@Mads I am still on the way with parallelism (new thing for me totally)... Looking at CPMP code i added maxtasksperchild=1 to yours to make sure each worker process one task. I got my spyder hanged.... and I added lines like print(\"start worker\"), so it hanged before any progress lines, which make me think it's loading. But maybe i made a mistake somewhere else.\nso how long load_dataset time is approx for you? I try to implement loading in the worker in the meantime</p>",
      "rawMarkdown": "Mads I am still on the way with parallelism (new thing for me totally)... Looking at CPMP code i added maxtasksperchild=1 to yours to make sure each worker process one task. I got my spyder hanged.... and I added lines like print(\"start worker\"), so it hanged before any progress lines, which make me think it's loading. But maybe i made a mistake somewhere else.\nso how long load_dataset time is approx for you? I try to implement loading in the worker in the meantime",
      "votes": null
    },
    {
      "id": "368187",
      "postDate": "08/09/2018 12:48:26",
      "content": "<p>@Blonde it hangs because the master is busy loading data for the workers.  Or you are working on Windows and inside notebooks, in which case you are in trouble :(</p>",
      "rawMarkdown": "Blonde it hangs because the master is busy loading data for the workers.  Or you are working on Windows and inside notebooks, in which case you are in trouble :(",
      "votes": null
    },
    {
      "id": "368221",
      "postDate": "08/09/2018 14:04:50",
      "content": "<p>BTW @Jack Vial, what IDE do you use on cloud? It's a bit out of topic, but related. </p>\n\n<p>I've no IT background at all (I googled how to use github :)), following simple instructions I installed anaconda, but could not start spyder .... then I installed pycharm using umake, but it would not start giving me Pycharm Startup Error: Unable to detect graphics environment. \nI found it: <a href=\"https://stackoverflow.com/questions/46124295/pycharm-startup-error-unable-to-detect-graphics-environment\">https://stackoverflow.com/questions/46124295/pycharm-startup-error-unable-to-detect-graphics-environment</a>\nI tried their solutions.\n1. I tried to install in opt but it tells me I do not have the privilege\n~$ /opt/pycharm/bin/pycharm.sh <br>\n2. Tried this one as well, did not work\n$ chmod +x pycharm.sh , \nrun pycharm with ./pycharm.sh <br>\n3. I install java which is not headless, but the same mistake.  :( I do not know what else to do/try, maybe some other IDE will work out better... \nPS I set up google cloud as it was easier then AWS, chose ubuntu 16.04 as in tutorial. Maybe there is a simple instruction that works for community pycharm, with empty ubunta (maybe I chose a wrong tutorial, i tied two with snap as well, with snap did not work)?   Maybe you could help me a bit with advice as you are the cloud user? I wanted to use cloud for other projects, so need to sort it not only for this competition</p>",
      "rawMarkdown": "BTW @Jack Vial, what IDE do you use on cloud? It's a bit out of topic, but related. \n\nI've no IT background at all (I googled how to use github :)), following simple instructions I installed anaconda, but could not start spyder .... then I installed pycharm using umake, but it would not start giving me Pycharm Startup Error: Unable to detect graphics environment. \nI found it: https://stackoverflow.com/questions/46124295/pycharm-startup-error-unable-to-detect-graphics-environment\nI tried their solutions.\n1. I tried to install in opt but it tells me I do not have the privilege\n~$ /opt/pycharm/bin/pycharm.sh  \n2. Tried this one as well, did not work\n$ chmod +x pycharm.sh , \nrun pycharm with ./pycharm.sh   \n3. I install java which is not headless, but the same mistake.  :( I do not know what else to do/try, maybe some other IDE will work out better... \nPS I set up google cloud as it was easier then AWS, chose ubuntu 16.04 as in tutorial. Maybe there is a simple instruction that works for community pycharm, with empty ubunta (maybe I chose a wrong tutorial, i tied two with snap as well, with snap did not work)?   Maybe you could help me a bit with advice as you are the cloud user? I wanted to use cloud for other projects, so need to sort it not only for this competition",
      "votes": null
    },
    {
      "id": "368237",
      "postDate": "08/09/2018 14:33:54",
      "content": "<p>@CPMP I implemented your code and it also hangs... I googled and found it is a spyder problem: <a href=\"https://github.com/spyder-ide/spyder/issues/2937\">https://github.com/spyder-ide/spyder/issues/2937</a>. \nWhat IDE do you use for successful parallel computing in python? Will jupyter work? Or better community pycharm?  Or better something else? \nI removed all other parameters, so it's not the issue. </p>\n\n<p>Also, this line is not a pseudo code, it's a valid line, correct? params = [i for i in range(10)]  I just never met such syntax before)</p>",
      "rawMarkdown": "CPMP I implemented your code and it also hangs... I googled and found it is a spyder problem: https://github.com/spyder-ide/spyder/issues/2937. \nWhat IDE do you use for successful parallel computing in python? Will jupyter work? Or better community pycharm?  Or better something else? \nI removed all other parameters, so it's not the issue. \n\nAlso, this line is not a pseudo code, it's a valid line, correct? params = [i for i in range(10)]  I just never met such syntax before)",
      "votes": null
    },
    {
      "id": "368239",
      "postDate": "08/09/2018 14:38:45",
      "content": "<p>I write and test my code locally in VSCode then push to github, then ssh into the AWS machine and pull the changes down from github. If I do need to make changes to the code directly on the AWS machine I use the Vim text editor which runs in the terminal  and is usually installed by default on Ubuntu and some other Linux distributions.</p>",
      "rawMarkdown": "I write and test my code locally in VSCode then push to github, then ssh into the AWS machine and pull the changes down from github. If I do need to make changes to the code directly on the AWS machine I use the Vim text editor which runs in the terminal  and is usually installed by default on Ubuntu and some other Linux distributions.",
      "votes": null
    },
    {
      "id": "368240",
      "postDate": "08/09/2018 14:39:26",
      "content": "<p>@CPMP yes, I am on Windows and inside spyder... I still cannot set up pycharm on ubuntu 16, used few tutorials and get Pycharm Startup Error: Unable to detect graphics environment, reinstalled, reinstalled java, googled that problem and tried all solutions I found, and still no progress...  I should rename my account to Dumb Blonde :))). What do you use for python on linux, is it ubunta? Can I do parallel computing in windows 7 or it's not gonna work? </p>",
      "rawMarkdown": "CPMP yes, I am on Windows and inside spyder... I still cannot set up pycharm on ubuntu 16, used few tutorials and get Pycharm Startup Error: Unable to detect graphics environment, reinstalled, reinstalled java, googled that problem and tried all solutions I found, and still no progress...  I should rename my account to Dumb Blonde :))). What do you use for python on linux, is it ubunta? Can I do parallel computing in windows 7 or it's not gonna work?",
      "votes": null
    },
    {
      "id": "368242",
      "postDate": "08/09/2018 14:48:46",
      "content": "<p>Yes, i have Vim. I still did not get it, when you take your ready code from github what program do you use to run it? you need some IDE, don't you? (I am not a linux user)</p>",
      "rawMarkdown": "Yes, i have Vim. I still did not get it, when you take your ready code from github what program do you use to run it? you need some IDE, don't you? (I am not a linux user)",
      "votes": null
    },
    {
      "id": "368244",
      "postDate": "08/09/2018 14:54:12",
      "content": "<p>I never managed to get Python parallelism work on Windows, but I was not as fluent with Python as I am now.</p>\n\n<p>On ubuntu why don't you use Jupyter Notebook?  It is quite convenient for data science and machine learning.  I use ubuntu 14.04, or ubuntu 16.04 on Intel machines, and RHEL on a Power machine.</p>",
      "rawMarkdown": "I never managed to get Python parallelism work on Windows, but I was not as fluent with Python as I am now.\n\nOn ubuntu why don't you use Jupyter Notebook?  It is quite convenient for data science and machine learning.  I use ubuntu 14.04, or ubuntu 16.04 on Intel machines, and RHEL on a Power machine.",
      "votes": null
    },
    {
      "id": "368245",
      "postDate": "08/09/2018 14:54:18",
      "content": "<p>I run the code from the terminal. Something like <code>python process_submission.py</code></p>",
      "rawMarkdown": "I run the code from the terminal. Something like `python process_submission.py`",
      "votes": null
    },
    {
      "id": "368246",
      "postDate": "08/09/2018 14:55:58",
      "content": "<p>Jack, How did you parallelize code?  </p>",
      "rawMarkdown": "Jack, How did you parallelize code?",
      "votes": null
    },
    {
      "id": "368252",
      "postDate": "08/09/2018 15:08:11",
      "content": "<p>@CPMP Thank you. I installed Jupyter Notebook on google cloud to learn fastai course. It starts a kernel but then kernel dead. \"The kernel has died, and the automatic restart has failed\". Found the story with the mistakes, troubleshooting... I think you can add to your discussion for newbies -- USE LINUX !!! </p>",
      "rawMarkdown": "CPMP Thank you. I installed Jupyter Notebook on google cloud to learn fastai course. It starts a kernel but then kernel dead. \"The kernel has died, and the automatic restart has failed\". Found the story with the mistakes, troubleshooting... I think you can add to your discussion for newbies -- USE LINUX !!!",
      "votes": null
    },
    {
      "id": "368253",
      "postDate": "08/09/2018 15:11:40",
      "content": "<p>ou, you can do just like that from a terminal, cool... but how it gets all the imports? I put a clusterer class in a separate py file and then do import from it, maybe it's not a good idea... and I use codes from trackml library and import from then load dataset and scoring. How do you handle that? Do you put all that in one single file? </p>",
      "rawMarkdown": "ou, you can do just like that from a terminal, cool... but how it gets all the imports? I put a clusterer class in a separate py file and then do import from it, maybe it's not a good idea... and I use codes from trackml library and import from then load dataset and scoring. How do you handle that? Do you put all that in one single file?",
      "votes": null
    },
    {
      "id": "368260",
      "postDate": "08/09/2018 15:18:49",
      "content": "<p>@CPMP I do something like this:</p>\n\n<pre><code>from trackml.dataset import load_event, load_dataset\nfrom trackml.score import score_event\nfrom joblib import Parallel, delayed\nimport multiprocessing\nimport time\nnum_cores = multiprocessing.cpu_count()\npath_to_test = \"./../data/test\"\n\ndef processOneSubmissionEvent(event_id, hits, cells):\n    model = Clusterer()\n    labels = model.run(hits)\n    submission = create_one_event_submission(event_id, hits, labels)\n    return (event_id, submission)\n\nprocessed_events = Parallel(n_jobs=num_cores)(delayed(processOneSubmissionEvent)(event_id, hits, cells) for event_id, hits, cells in load_dataset(path_to_test, parts=['hits', 'cells']))\n    test_dataset_submissions = []\n    for event in processed_events:\n        test_dataset_submissions.append(event[1])\n\n    # Create submission file\n    ts = str(time.time()).split(\".\")[0]\n    final_submussion = pd.concat(test_dataset_submissions, axis=0)\n    final_submussion.to_csv(\"./../submissions/\" + ts + \"_submission.csv\", index=False)\n    print(\"submission created\")\n</code></pre>",
      "rawMarkdown": "CPMP I do something like this:\n\n    from trackml.dataset import load_event, load_dataset\n    from trackml.score import score_event\n    from joblib import Parallel, delayed\n    import multiprocessing\n    import time\n    num_cores = multiprocessing.cpu_count()\n    path_to_test = \"./../data/test\"\n    \n    def processOneSubmissionEvent(event_id, hits, cells):\n        model = Clusterer()\n        labels = model.run(hits)\n        submission = create_one_event_submission(event_id, hits, labels)\n        return (event_id, submission)\n    \n    processed_events = Parallel(n_jobs=num_cores)(delayed(processOneSubmissionEvent)(event_id, hits, cells) for event_id, hits, cells in load_dataset(path_to_test, parts=['hits', 'cells']))\n        test_dataset_submissions = []\n        for event in processed_events:\n            test_dataset_submissions.append(event[1])\n            \n        # Create submission file\n        ts = str(time.time()).split(\".\")[0]\n        final_submussion = pd.concat(test_dataset_submissions, axis=0)\n        final_submussion.to_csv(\"./../submissions/\" + ts + \"_submission.csv\", index=False)\n        print(\"submission created\")",
      "votes": null
    },
    {
      "id": "368261",
      "postDate": "08/09/2018 15:19:17",
      "content": "<p>@CPMP Load time from disk for all the data is 30 seconds for me. 27 seconds after it is cached (I'm on Ubuntu). Lets say the data is loaded and most of the 27 seconds is making dataframes. It has to pass these dataframes to the workers, so we have another memory copy. Lets just say it takes the same amount of time and double the data time. That is still less than 1 sec per event. My model runs a little slower than 1 sec per event.</p>\n\n<p>@Blonde You can see my load time above. Also load_dataset is a generator, it only loads or does something when you try to get something from it.\nif you do something like:</p>\n\n<pre><code>a=load_dataset(path_to_test, parts=['hits', 'cells'])\n</code></pre>\n\n<p>nothing happens, it just gets thing ready to use.\nI really see no need in setting maxtasksperchild. It could make things slower if you had the same data that was used in each child. I would use it if the event data was dramatically different in size per event and memory was an issue, sometimes the actual python task doesn't like to let go of ram or some lib has a memory leak.</p>\n\n<p>I've ran this code successfully in notebook and as a script. I'm running 10 threads on 6 core i7, there is about a 30 second pause before the works start churning and they all seem to start working at once. So my calculations above could be off by a second and data load is take 3 seconds per event. With 10 threads I'm using less than 2GB of ram.</p>",
      "rawMarkdown": "CPMP Load time from disk for all the data is 30 seconds for me. 27 seconds after it is cached (I'm on Ubuntu). Lets say the data is loaded and most of the 27 seconds is making dataframes. It has to pass these dataframes to the workers, so we have another memory copy. Lets just say it takes the same amount of time and double the data time. That is still less than 1 sec per event. My model runs a little slower than 1 sec per event.\n\n@Blonde You can see my load time above. Also load_dataset is a generator, it only loads or does something when you try to get something from it.\nif you do something like:\n\n    a=load_dataset(path_to_test, parts=['hits', 'cells'])\n\nnothing happens, it just gets thing ready to use.\nI really see no need in setting maxtasksperchild. It could make things slower if you had the same data that was used in each child. I would use it if the event data was dramatically different in size per event and memory was an issue, sometimes the actual python task doesn't like to let go of ram or some lib has a memory leak.\n \nI've ran this code successfully in notebook and as a script. I'm running 10 threads on 6 core i7, there is about a 30 second pause before the works start churning and they all seem to start working at once. So my calculations above could be off by a second and data load is take 3 seconds per event. With 10 threads I'm using less than 2GB of ram.",
      "votes": null
    },
    {
      "id": "368262",
      "postDate": "08/09/2018 15:20:32",
      "content": "<p><a href=\"/blonde\">@blonde</a> The imports will be imported using the import statements usually at the top of the script e.g.</p>\n\n<p>from trackml.dataset import load_event, load_dataset</p>",
      "rawMarkdown": "blonde The imports will be imported using the import statements usually at the top of the script e.g.\n\nfrom trackml.dataset import load_event, load_dataset",
      "votes": null
    },
    {
      "id": "368265",
      "postDate": "08/09/2018 15:28:16",
      "content": "<p>ok, got it, will try. Do I need to put all those progs to a special folder on server where my python exe file is, or it can run from any location? Sorry for this dumb questions, i am totally new to linux, cloud, python, ML... only google and people save me</p>",
      "rawMarkdown": "ok, got it, will try. Do I need to put all those progs to a special folder on server where my python exe file is, or it can run from any location? Sorry for this dumb questions, i am totally new to linux, cloud, python, ML... only google and people save me",
      "votes": null
    },
    {
      "id": "368284",
      "postDate": "08/09/2018 16:09:12",
      "content": "<p>Thank you @Mads. I use your pool.starmap(worker,load_dataset(path_to_test, parts=['hits', 'cells'])), but i save submission for each event individually in worker. My problem is Windows 7. Now struggling to set up jupyter on linux cloud... installed but kernel dead :(, doing troubleshooting, will see ... Now I see that kaggle is for linux users :)</p>",
      "rawMarkdown": "Thank you @Mads. I use your pool.starmap(worker,load_dataset(path_to_test, parts=['hits', 'cells'])), but i save submission for each event individually in worker. My problem is Windows 7. Now struggling to set up jupyter on linux cloud... installed but kernel dead :(, doing troubleshooting, will see ... Now I see that kaggle is for linux users :)",
      "votes": null
    },
    {
      "id": "368337",
      "postDate": "08/09/2018 18:14:02",
      "content": "<p>Our current solution (0.646) runs in R and takes about 150 minutes to create the testset submission using lots of multiprocessing techniques :-)</p>",
      "rawMarkdown": "Our current solution (0.646) runs in R and takes about 150 minutes to create the testset submission using lots of multiprocessing techniques :-)",
      "votes": null
    },
    {
      "id": "368339",
      "postDate": "08/09/2018 18:16:02",
      "content": "<p>&gt; It does not make much difference till the number of remaining events is less than n_proc.</p>\n\n<p>I use parallelization inside one event, every dbscan is a task. So one task is less than 1 second. I think it will be not efficient to kill processes.\nBy the way, don't you use parallelization in one event? So how do you test weigths etc? </p>",
      "rawMarkdown": "&gt; It does not make much difference till the number of remaining events is less than n_proc.\n\nI use parallelization inside one event, every dbscan is a task. So one task is less than 1 second. I think it will be not efficient to kill processes.\nBy the way, don't you use parallelization in one event? So how do you test weigths etc?",
      "votes": null
    },
    {
      "id": "368341",
      "postDate": "08/09/2018 18:21:38",
      "content": "<p>Check out Docker running a linux/data science container on windows. It might be easier. I haven't set up  personally for this, but on other things on a mac.</p>",
      "rawMarkdown": "Check out Docker running a linux/data science container on windows. It might be easier. I haven't set up  personally for this, but on other things on a mac.",
      "votes": null
    },
    {
      "id": "368360",
      "postDate": "08/09/2018 19:04:50",
      "content": "<p>R is better (maybe for multiprocessing  )? Why not python? </p>\n\n<p>It takes 2 or 3 days for me to create the testset submission. :((</p>",
      "rawMarkdown": "R is better (maybe for multiprocessing  )? Why not python? \n\nIt takes 2 or 3 days for me to create the testset submission. :((",
      "votes": null
    },
    {
      "id": "368376",
      "postDate": "08/09/2018 19:36:10",
      "content": "<p>Sergey, same for me 3 days for the miserable score on a usual windows laptop (cause I did not know about linux and that I can parallel...). I still struggle to set python IDE on linux, but I keep trying... now uninstalled everything reinstalled... and of course, nothing works :) This ML problem turned our to be linux problem for me :))) but then i can enjoy running progs on linux afterwards (if i set it up)</p>",
      "rawMarkdown": "Sergey, same for me 3 days for the miserable score on a usual windows laptop (cause I did not know about linux and that I can parallel...). I still struggle to set python IDE on linux, but I keep trying... now uninstalled everything reinstalled... and of course, nothing works :) This ML problem turned our to be linux problem for me :))) but then i can enjoy running progs on linux afterwards (if i set it up)",
      "votes": null
    },
    {
      "id": "368395",
      "postDate": "08/09/2018 20:11:32",
      "content": "<blockquote>\n  <p>My model runs a little slower than 1 sec per event.</p>\n</blockquote>\n\n<p>@Mads,  you don't seem to have a clear view of how parallel computing works here.  First, you are off by a factor of 40 or so on the overhead for data loading.  indeed, here is how things run:</p>\n\n<ol>\n<li><p>The master loads all data.  This takes 27 seconds.</p></li>\n<li><p>The master dispatches the data to workers.  It does it one worker at a time because of the GIL.  You say this takes about the same time, which means that on average workers get their data in 27/2 seconds.  As a result, workers wait 3/2 * 27 seconds on average, i.e. 40 seconds, not one second.</p></li>\n</ol>\n\n<p>Granted, this is not a big deal if workers runs for hours.</p>\n\n<p>But if workers run for hours then you have another issue.  With default settings, the Pool assigns more than one tasks per worker.  With 125 tasks and 20 workers (my case), it assigns 2 tasks per worker.  This means that after processing the first 120 events, it assigns the last 5 events to 3 workers, not 5.  It means that processing the last 5 events takes twice the time to process one event, even if you have more than 5 workers.  It my case it means that the overall computation takes about 8 times the time it takes to process one event, instead of 7. it means 10 extra hours.  Not negligible at all.</p>\n\n<p>This said, if you think your way is fine, great.  Do as you wish.  But you posted your (bad) way of doing things, which is why I put the dots on the i.</p>",
      "rawMarkdown": "&gt; My model runs a little slower than 1 sec per event.\n\n@Mads,  you don't seem to have a clear view of how parallel computing works here.  First, you are off by a factor of 40 or so on the overhead for data loading.  indeed, here is how things run:\n\n 1. The master loads all data.  This takes 27 seconds.\n\n 2.  The master dispatches the data to workers.  It does it one worker at a time because of the GIL.  You say this takes about the same time, which means that on average workers get their data in 27/2 seconds.  As a result, workers wait 3/2 * 27 seconds on average, i.e. 40 seconds, not one second.\n\nGranted, this is not a big deal if workers runs for hours.\n\nBut if workers run for hours then you have another issue.  With default settings, the Pool assigns more than one tasks per worker.  With 125 tasks and 20 workers (my case), it assigns 2 tasks per worker.  This means that after processing the first 120 events, it assigns the last 5 events to 3 workers, not 5.  It means that processing the last 5 events takes twice the time to process one event, even if you have more than 5 workers.  It my case it means that the overall computation takes about 8 times the time it takes to process one event, instead of 7. it means 10 extra hours.  Not negligible at all.\n\nThis said, if you think your way is fine, great.  Do as you wish.  But you posted your (bad) way of doing things, which is why I put the dots on the i.",
      "votes": null
    },
    {
      "id": "368417",
      "postDate": "08/09/2018 21:19:09",
      "content": "<p>ok, while I unsuccessfully try to install everything on google cloud i drove to Uni where they have linux. Now both variants start, but for @Mads version i have:\n'Pool' object has no attribute 'starmap'\nand for @CPMP version I have: 'total' is an invalid keyword argument for this function (for ls   = pool.map(work_sub, params, chunksize=1) )\nI think it's because there is python 2.7 there and maybe older multiprocessing and Pool... but i cannot update it as it's not my server... I keep optimizing my clustering, hopefully I'll be able to sort it out and run a final submission </p>",
      "rawMarkdown": "ok, while I unsuccessfully try to install everything on google cloud i drove to Uni where they have linux. Now both variants start, but for @Mads version i have:\n'Pool' object has no attribute 'starmap'\nand for @CPMP version I have: 'total' is an invalid keyword argument for this function (for ls   = pool.map(work_sub, params, chunksize=1) )\nI think it's because there is python 2.7 there and maybe older multiprocessing and Pool... but i cannot update it as it's not my server... I keep optimizing my clustering, hopefully I'll be able to sort it out and run a final submission",
      "votes": null
    },
    {
      "id": "368419",
      "postDate": "08/09/2018 21:23:23",
      "content": "<p>@CPMP If I have time I will do the load in the worker and time each run. Like I said at the end post it takes about 30 seconds for things to get going. After that I don't see any system load decrease until things taper off at the end. And it is clear each process is running only one event.</p>\n\n<blockquote>\n  <p>maxtasksperchild is the number of tasks a worker process can complete before it will exit and be replaced with a fresh worker process, to enable unused resources to be freed. The default maxtasksperchild is None, which means worker processes will live as long as the pool.</p>\n</blockquote>\n\n<p>This has nothing to do with \"chunksize.\" I think you might be confusing the two.</p>\n\n<blockquote>\n  <p>map(func, iterable[, chunksize])\n  A parallel equivalent of the map() built-in function (it supports only one iterable argument though). It blocks until the result is ready.</p>\n  \n  <p>This method chops the iterable into a number of chunks which it submits to the process pool as separate tasks. The (approximate) size of these chunks can be specified by setting chunksize to a positive integer.</p>\n</blockquote>\n\n<p>And your math doesn't seem to match my results. If it was 40 seconds per task, it would take 400 seconds for things to start working for 10 tasks, it takes 30 seconds as I mentioned. </p>",
      "rawMarkdown": "CPMP If I have time I will do the load in the worker and time each run. Like I said at the end post it takes about 30 seconds for things to get going. After that I don't see any system load decrease until things taper off at the end. And it is clear each process is running only one event.\n\n &gt;maxtasksperchild is the number of tasks a worker process can complete before it will exit and be replaced with a fresh worker process, to enable unused resources to be freed. The default maxtasksperchild is None, which means worker processes will live as long as the pool.\n\nThis has nothing to do with \"chunksize.\" I think you might be confusing the two.\n\n&gt;map(func, iterable[, chunksize])\n&gt;A parallel equivalent of the map() built-in function (it supports only one iterable argument though). It blocks until the result is ready.\n\n&gt;This method chops the iterable into a number of chunks which it submits to the process pool as separate tasks. The (approximate) size of these chunks can be specified by setting chunksize to a positive integer.\n\nAnd your math doesn't seem to match my results. If it was 40 seconds per task, it would take 400 seconds for things to start working for 10 tasks, it takes 30 seconds as I mentioned.",
      "votes": null
    },
    {
      "id": "368420",
      "postDate": "08/09/2018 21:27:43",
      "content": "<p>@Blonde starmap is in python3.3 and on. You could see if python3 is on the system by running your script with python3 your_script.py\nIf that doesn't do it there is always setting up a virtual_env and having python3 in that.</p>",
      "rawMarkdown": "Blonde starmap is in python3.3 and on. You could see if python3 is on the system by running your script with python3 your_script.py\nIf that doesn't do it there is always setting up a virtual_env and having python3 in that.",
      "votes": null
    },
    {
      "id": "368422",
      "postDate": "08/09/2018 21:33:55",
      "content": "<p>R is not better than python for multiprocessing.\nI just figured out that my algorithm, for some reason that I don't know, runs faster in R :-)</p>",
      "rawMarkdown": "R is not better than python for multiprocessing.\nI just figured out that my algorithm, for some reason that I don't know, runs faster in R :-)",
      "votes": null
    },
    {
      "id": "368494",
      "postDate": "08/10/2018 02:57:45",
      "content": "<blockquote>\n  <p>runs faster in R :-)</p>\n</blockquote>\n\n<p>Wow! ;)</p>",
      "rawMarkdown": "&gt; runs faster in R :-)\n\nWow! ;)",
      "votes": null
    },
    {
      "id": "368497",
      "postDate": "08/10/2018 03:01:04",
      "content": "<blockquote>\n  <p>If it was 40 seconds per task, it would take 400 seconds for things to start working for 10 tasks,</p>\n</blockquote>\n\n<p>You should read more carefully.  I wrote that each of your worker would start roughly 40 second after the master starts.  </p>\n\n<p>Anyway, \"on ne fait pas boir un ane qui n'a pas soif\".  Sorry for the French, I don't know the English equivalent.  </p>\n\n<p>As I said, if you're fine with your code, fine with me.  I know for sure, because I measured it, that my way saves me 10 hours per submission. </p>",
      "rawMarkdown": "&gt; If it was 40 seconds per task, it would take 400 seconds for things to start working for 10 tasks,\n\nYou should read more carefully.  I wrote that each of your worker would start roughly 40 second after the master starts.  \n\nAnyway, \"on ne fait pas boir un ane qui n'a pas soif\".  Sorry for the French, I don't know the English equivalent.  \n\nAs I said, if you're fine with your code, fine with me.  I know for sure, because I measured it, that my way saves me 10 hours per submission.",
      "votes": null
    },
    {
      "id": "368498",
      "postDate": "08/10/2018 03:04:05",
      "content": "<p>I use parallelism for events and I don't use parallelism inside each event.  It is a coarse grain parallelism.  You use fine grain parallelism, which makes the master a bottleneck very easily.   My gut feeling is that this is not going to be as efficient as coarse grain parallelism, and it probably explain why it does not scale beyond 3 parallel processes.</p>",
      "rawMarkdown": "I use parallelism for events and I don't use parallelism inside each event.  It is a coarse grain parallelism.  You use fine grain parallelism, which makes the master a bottleneck very easily.   My gut feeling is that this is not going to be as efficient as coarse grain parallelism, and it probably explain why it does not scale beyond 3 parallel processes.",
      "votes": null
    },
    {
      "id": "368731",
      "postDate": "08/10/2018 16:23:00",
      "content": "<p>@CPMP I managed to parallel in Windows! It works in cmd, the trick is you need to add (otherwise, it hangs): \nif <strong>name</strong> == '<strong>main</strong>':\n    if use_parallel: \n        print('start worker')\n        pool = Pool(processes=n_proc, maxtasksperchild=1)\n        ls   = pool.map( work_sub, params, chunksize=1 )\n        pool.close()\n    else:\n        ls = [work_sub(param) for param in params]</p>\n\n<p>I never considered using just cmd, now I see it's better in some ways. Thank you for your code, and thanks to @Mads and  @Jack Vial for comments. Next step is doing it on linux... Maybe it would not help me to improve before the deadline, but I started kaggle to learn things and it was great learning fun! I cannot add emotion here, but you can imagine a little penguin dancing :))) </p>\n\n<p>PS. @CPMP I really appreciate this post! The only little thing: such obvious things for you and other experienced users are nice to publish a bit earlier, like a month before the deadline, for the complete beginners, who take time to make it and did not hear about parallel programming</p>",
      "rawMarkdown": "CPMP I managed to parallel in Windows! It works in cmd, the trick is you need to add (otherwise, it hangs): \nif __name__ == '__main__':\n    if use_parallel: \n        print('start worker')\n        pool = Pool(processes=n_proc, maxtasksperchild=1)\n        ls   = pool.map( work_sub, params, chunksize=1 )\n        pool.close()\n    else:\n        ls = [work_sub(param) for param in params]\n\nI never considered using just cmd, now I see it's better in some ways. Thank you for your code, and thanks to @Mads and  @Jack Vial for comments. Next step is doing it on linux... Maybe it would not help me to improve before the deadline, but I started kaggle to learn things and it was great learning fun! I cannot add emotion here, but you can imagine a little penguin dancing :))) \n\nPS. @CPMP I really appreciate this post! The only little thing: such obvious things for you and other experienced users are nice to publish a bit earlier, like a month before the deadline, for the complete beginners, who take time to make it and did not hear about parallel programming",
      "votes": null
    },
    {
      "id": "368745",
      "postDate": "08/10/2018 17:04:17",
      "content": "<p><a href=\"/blonde\">@blonde</a> </p>\n\n<blockquote>\n  <p>Do I need to put all those progs to a special folder on server where\n  my python exe file is</p>\n</blockquote>\n\n<p>If you are refereeing to the module imports then python will know where to look for them if they are installed with pip or anaconda. If the module you are importing is some file in your project then it needs to be in the same directory as the file that is importing it or you need to add it to your path so python knows where to look for it.</p>\n\n<blockquote>\n  <p>Sorry for this dumb questions, i am totally new to linux, cloud,\n  python, ML... only google and people save me</p>\n</blockquote>\n\n<p>There are no dumb questions! :)</p>",
      "rawMarkdown": "blonde \n\n&gt; Do I need to put all those progs to a special folder on server where\n&gt; my python exe file is\n\nIf you are refereeing to the module imports then python will know where to look for them if they are installed with pip or anaconda. If the module you are importing is some file in your project then it needs to be in the same directory as the file that is importing it or you need to add it to your path so python knows where to look for it.\n\n&gt; Sorry for this dumb questions, i am totally new to linux, cloud,\n&gt; python, ML... only google and people save me\n\nThere are no dumb questions! :)",
      "votes": null
    },
    {
      "id": "368758",
      "postDate": "08/10/2018 17:46:01",
      "content": "<p>@Jack Vial </p>\n\n<p>I managed to get my first paralleled project in Windows! With all my imports etc. \nWith no IT background I never considered using just cmd, now I see it's better in some ways. Thank you for your code and comments! </p>\n\n<p>The next step is doing it on linux... Maybe it would not help me to improve before the deadline, but I started kaggle to learn things and that was great learning fun! I cannot add emotion here, but you can imagine a little penguin dancing :)))</p>",
      "rawMarkdown": "Jack Vial \n\nI managed to get my first paralleled project in Windows! With all my imports etc. \nWith no IT background I never considered using just cmd, now I see it's better in some ways. Thank you for your code and comments! \n\nThe next step is doing it on linux... Maybe it would not help me to improve before the deadline, but I started kaggle to learn things and that was great learning fun! I cannot add emotion here, but you can imagine a little penguin dancing :)))",
      "votes": null
    },
    {
      "id": "368762",
      "postDate": "08/10/2018 17:57:57",
      "content": "<p>&gt; I managed to parallel in Windows! It works in cmd, the trick is you need to add (otherwise, it hangs): if name == 'main': </p>\n\n<p>Yeah, you can look at my kernel with parallelization (not events, but dbscans). I use Windows too.</p>\n\n<p><a href=\"https://www.kaggle.com/sergeyzlobin/unrolling-helices-baseline-python\">https://www.kaggle.com/sergeyzlobin/unrolling-helices-baseline-python</a></p>",
      "rawMarkdown": "&gt; I managed to parallel in Windows! It works in cmd, the trick is you need to add (otherwise, it hangs): if name == 'main': \n\nYeah, you can look at my kernel with parallelization (not events, but dbscans). I use Windows too.\n\nhttps://www.kaggle.com/sergeyzlobin/unrolling-helices-baseline-python",
      "votes": null
    },
    {
      "id": "368765",
      "postDate": "08/10/2018 18:05:13",
      "content": "<blockquote>\n  <p>My gut feeling is that this is not going to be as efficient as coarse grain parallelism, and it probably explain why it does not scale beyond 3 parallel processes.</p>\n</blockquote>\n\n<p>Yes, probably you're right. Anyway I don't have machines with more than 4 cores. </p>",
      "rawMarkdown": "&gt;  My gut feeling is that this is not going to be as efficient as coarse grain parallelism, and it probably explain why it does not scale beyond 3 parallel processes.\n\nYes, probably you're right. Anyway I don't have machines with more than 4 cores.",
      "votes": null
    },
    {
      "id": "368768",
      "postDate": "08/10/2018 18:27:36",
      "content": "<p><a href=\"/blonde\">@blonde</a> Congrats! Yes the cmd/terminal/shell is very useful and powerful, especially on linux. </p>\n\n<p>Definitely if not in this competition everything you learn will help in the next! Maybe you will need to put in a request for emojis! :)</p>",
      "rawMarkdown": "blonde Congrats! Yes the cmd/terminal/shell is very useful and powerful, especially on linux. \n\nDefinitely if not in this competition everything you learn will help in the next! Maybe you will need to put in a request for emojis! :)",
      "votes": null
    },
    {
      "id": "368771",
      "postDate": "08/10/2018 18:55:51",
      "content": "<p>I just prefer to use windows subsystem for linux. </p>",
      "rawMarkdown": "I just prefer to use windows subsystem for linux.",
      "votes": null
    },
    {
      "id": "368801",
      "postDate": "08/10/2018 21:04:54",
      "content": "<blockquote>\n  <p>PS. @CPMP I really appreciate this post! The only little thing: such obvious things for you and other experienced users are nice to publish a bit earlier, like a month before the deadline, for the complete beginners, who take time to make it and did not hear about parallel programming</p>\n</blockquote>\n\n<p>Better late than never ;)</p>",
      "rawMarkdown": "&gt; PS. @CPMP I really appreciate this post! The only little thing: such obvious things for you and other experienced users are nice to publish a bit earlier, like a month before the deadline, for the complete beginners, who take time to make it and did not hear about parallel programming\n\nBetter late than never ;)",
      "votes": null
    },
    {
      "id": "368949",
      "postDate": "08/11/2018 11:27:52",
      "content": "<p>@CPMP Better late than never ;)  True...  I now managed to start it on google cloud!!! 24 cpu (I was stupid and chose a wrong region instead of usa), so now running on linux...it should help me better. I thought it would be 24 times faster than laptop (yes, I did not prallel my calculations before, at all... ), it does not scale like that, it's still two times slower than laptop event time*124/24, but still...  Can't add emotion, but now you can imagine 24 little penguins dancing :))</p>\n\n<p>@Sergey, I saw your kernel a while ago but was unable to understand it (I've zero IT background), so now as I got my better submissions running I was planning exactly that: parallel in the main to optimise things better... so much more I could have done if I knew about parallel... optimising things much much faster and for many different features, and on 10 events, instead of picking one, waiting for 30 min for a single calculation and hoping for the best :)... Anyway, with my zero IT/linux background bronze is ok for the debut.  And now I feel I'll be able to stay there.</p>",
      "rawMarkdown": "CPMP Better late than never ;)  True...  I now managed to start it on google cloud!!! 24 cpu (I was stupid and chose a wrong region instead of usa), so now running on linux...it should help me better. I thought it would be 24 times faster than laptop (yes, I did not prallel my calculations before, at all... ), it does not scale like that, it's still two times slower than laptop event time*124/24, but still...  Can't add emotion, but now you can imagine 24 little penguins dancing :))\n\n@Sergey, I saw your kernel a while ago but was unable to understand it (I've zero IT background), so now as I got my better submissions running I was planning exactly that: parallel in the main to optimise things better... so much more I could have done if I knew about parallel... optimising things much much faster and for many different features, and on 10 events, instead of picking one, waiting for 30 min for a single calculation and hoping for the best :)... Anyway, with my zero IT/linux background bronze is ok for the debut.  And now I feel I'll be able to stay there.",
      "votes": null
    },
    {
      "id": "368951",
      "postDate": "08/11/2018 11:29:15",
      "content": "<p>@Jack Vial </p>\n\n<p>I now managed to start it on google cloud!!! 24 cpu, it does not scale like that, it's still two times slower than laptop event time*124/24, but still... I'll ask for the emotions,  but for now you can imagine 24 little penguins dancing :))</p>",
      "rawMarkdown": "Jack Vial \n\nI now managed to start it on google cloud!!! 24 cpu, it does not scale like that, it's still two times slower than laptop event time*124/24, but still... I'll ask for the emotions,  but for now you can imagine 24 little penguins dancing :))",
      "votes": null
    },
    {
      "id": "368969",
      "postDate": "08/11/2018 12:47:35",
      "content": "<blockquote>\n  <p>I don't have machines with more than 4 cores. </p>\n</blockquote>\n\n<p>Yes, but you're using only 3.  If you could use 4 then you would get a 33% speedup ;)</p>",
      "rawMarkdown": "&gt; I don't have machines with more than 4 cores. \n\nYes, but you're using only 3.  If you could use 4 then you would get a 33% speedup ;)",
      "votes": null
    },
    {
      "id": "369118",
      "postDate": "08/12/2018 02:49:41",
      "content": "<p>When using multiple processes, Python copies global variables to each worker and doesn't sync any changes made to these global variables back to the master process since all the processes are concurrent. Each process could merge the current track and a global variable containing the best track, but each process wouldn't be able to update the global variable containing the best track. How do you make it work?</p>",
      "rawMarkdown": "When using multiple processes, Python copies global variables to each worker and doesn't sync any changes made to these global variables back to the master process since all the processes are concurrent. Each process could merge the current track and a global variable containing the best track, but each process wouldn't be able to update the global variable containing the best track. How do you make it work?",
      "votes": null
    },
    {
      "id": "369153",
      "postDate": "08/12/2018 06:16:11",
      "content": "<p>You can use Manger list from multiprocessing. This should help you manage global variables across multi processes.</p>\n\n<p><code>from multiprocessing import Manager</code></p>",
      "rawMarkdown": "You can use Manger list from multiprocessing. This should help you manage global variables across multi processes.\n\n`from multiprocessing import Manager`",
      "votes": null
    },
    {
      "id": "369186",
      "postDate": "08/12/2018 08:51:28",
      "content": "<p>You can use the <code>work_sub</code> return value for passing results back to the master.  The list <code>ls</code> will contain all the returned values.  Or you can have the master code read the saved event submissions.  This is what I do after the parallel loop is executed.</p>",
      "rawMarkdown": "You can use the `work_sub` return value for passing results back to the master.  The list `ls` will contain all the returned values.  Or you can have the master code read the saved event submissions.  This is what I do after the parallel loop is executed.",
      "votes": null
    },
    {
      "id": "369239",
      "postDate": "08/12/2018 13:39:46",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> </p>\n\n<blockquote>\n  <p>you can have the master code read the saved event submissions</p>\n</blockquote>\n\n<p>This is what I currently do; Whats different is that I parallelize over each DBSCAN iteration while the code above parallelizes over each event. I think that if I'm using this approach, using multiprocessing.manager should solve the problem. I guess that if you're running each event on a different event, your DBSCAN n_jobs value is 1. Would it be a good idea to use part of the cores to parallelize over the events and part of the cores to parallelize over the DBSCAN iterations?</p>",
      "rawMarkdown": "cpmpml \n\n&gt; you can have the master code read the saved event submissions\n\nThis is what I currently do; Whats different is that I parallelize over each DBSCAN iteration while the code above parallelizes over each event. I think that if I'm using this approach, using multiprocessing.manager should solve the problem. I guess that if you're running each event on a different event, your DBSCAN n_jobs value is 1. Would it be a good idea to use part of the cores to parallelize over the events and part of the cores to parallelize over the DBSCAN iterations?",
      "votes": null
    },
    {
      "id": "369245",
      "postDate": "08/12/2018 13:49:06",
      "content": "<p>TL;DR Do you see a linear speedup as you increase the number of workers in you case?  I see linear speedup for 40 cores.</p>\n\n<p>Parallelism has some overhead, and the coarser the task, the better.  It is why I use parallelism per event and computation for one event is single threaded.  This way scales almost perfectly for 40 cores.   I don't see why one would need more parallelism unless one has more cores than events...</p>\n\n<p>Finer grained parallelism is better handle via multi threading, but Python does not support multi threading.  Sure, there is a package for it, but in general only one thread can execute at a time because of the GIL.</p>",
      "rawMarkdown": "TL;DR Do you see a linear speedup as you increase the number of workers in you case?  I see linear speedup for 40 cores.\n\nParallelism has some overhead, and the coarser the task, the better.  It is why I use parallelism per event and computation for one event is single threaded.  This way scales almost perfectly for 40 cores.   I don't see why one would need more parallelism unless one has more cores than events...\n\nFiner grained parallelism is better handle via multi threading, but Python does not support multi threading.  Sure, there is a package for it, but in general only one thread can execute at a time because of the GIL.",
      "votes": null
    },
    {
      "id": "369292",
      "postDate": "08/12/2018 16:19:52",
      "content": "<p>I used \"preemptable\" server instances from Google Cloud (to save money...).  Since you can expect to be preempted a couple of times per day, I chose to use parallelism to finish one event as quickly as possible, even if it isn't the most efficient method.  Writing code to recover one event from an interruption seemed easier than code to recover 90+ ... but maybe I was just being lazy!</p>\n\n<p>My local code was written to process one event per thread, but as my solution complexity grew where I couldn't generate a submission in a couple of days on a 10-core CPU, I moved to the cloud.</p>",
      "rawMarkdown": "I used \"preemptable\" server instances from Google Cloud (to save money...).  Since you can expect to be preempted a couple of times per day, I chose to use parallelism to finish one event as quickly as possible, even if it isn't the most efficient method.  Writing code to recover one event from an interruption seemed easier than code to recover 90+ ... but maybe I was just being lazy!\n\nMy local code was written to process one event per thread, but as my solution complexity grew where I couldn't generate a submission in a couple of days on a 10-core CPU, I moved to the cloud.",
      "votes": null
    },
    {
      "id": "369295",
      "postDate": "08/12/2018 16:28:12",
      "content": "<p>John, this is indeed a good reason to use parallelism within one event computation.</p>",
      "rawMarkdown": "John, this is indeed a good reason to use parallelism within one event computation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 367780,
      "author_name": "sergeyzlobin",
      "author_url": "",
      "post_date": "08/08/2018 14:08:15",
      "content": "<blockquote>\n  <p>passing data means the master process must load it first, and this can become a real bottleneck for the whole computation.</p>\n</blockquote>\n\n<p>Do you know a time for passing data? Suppose 1 dbscan is about 1 second. Is it comparable to it? I pass 'hits' values and other params. I thought passing is fast. :(</p>",
      "votes": null,
      "replies": [
        {
          "id": 367783,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/08/2018 14:22:49",
          "content": "<p>The problem is not that passing is fast or not.  The problem is that python cannot execute more than one thread at a time because of the GIL.  It means that the master can pass data to only one worker at a time, making it non parallel...  The effect is worse when your worker runs fast, or if you have many workers.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 367821,
          "author_name": "sergeyzlobin",
          "author_url": "",
          "post_date": "08/08/2018 15:52:00",
          "content": "<p>I've taken measurements in my case. The overhead of parallelization is about 8% (including parameter passing). A lot, but not deadly.</p>\n\n<p>And thanks for this info! I didn't think about this overhead.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 367843,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/08/2018 16:31:29",
          "content": "<p>Parallelism overhead is so small it is not measurable with my code ;)  I agree that 8% is not deadly, I guess you don't use many parallel processes.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 367874,
          "author_name": "sergeyzlobin",
          "author_url": "",
          "post_date": "08/08/2018 18:13:50",
          "content": "<p>&gt; I guess you don't use many parallel processes. </p>\n\n<p>You're right. I use 3 processes (sometimes on different machines). If use all 4 cores, all becomes slow and I can't do anything else.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 367876,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/08/2018 18:15:44",
          "content": "<p>See, your way is not enabling you to use your 4 cores...  I bet you would be able to use them if you move to the way I describe above.</p>\n\n<p>Your overhead is way higher than you think, you should measure it when running with 4 workers (one per core).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 367790,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/08/2018 14:42:54",
      "content": "<p>The worst I can imagine is to load all events data in the master, then pass it to each worker.  Then the worker code selects the relevant part of it.  </p>\n\n<p>With this pattern you load the 125 events data <code>n_proc + 1</code> times, overloading your machine memory, and slowing it dramatically.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 367800,
      "author_name": "returnofsputnik",
      "author_url": "",
      "post_date": "08/08/2018 15:15:03",
      "content": "<p>\"You can pass additional parameters but be careful, all data you pass there can create a bottleneck for the whole computation, stick to simple parameters like numerical values. \" - Ok, but then if you are required to get the data for event i, you are still required to access the event data by doing something like <code>df.loc[df['event']==i]</code>. So you are forced to pass in <code>df</code>, unless you make <code>df</code> a global variable</p>",
      "votes": null,
      "replies": [
        {
          "id": 367811,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/08/2018 15:36:31",
          "content": "<p>You are describing one of the worst, if not the worst, way of implementing parallelism: you make all event data copied in each worker.</p>\n\n<blockquote>\n  <p>you are still required to access the event data by doing something like df.loc[df['event']==i]</p>\n</blockquote>\n\n<p>Absolutely not, you can load each event data independently in each worker.  For instance, here is the start of my worker code:</p>\n\n<pre><code>def get_event(i):\n    prefix = 'event000000'\n    if i &lt; 100:\n        prefix = prefix +'0'\n    if i &lt; 10:\n        prefix = prefix +'0'\n    return prefix + str(i)\n\nbase_path = '/home/jfpuget/Kaggle/TrackML/'\n\ndef work_sub(param):\n    (i, &lt;other parameters&gt;) = param\n\n    event = get_event(i)\n    print('event:', event)\n    hits = pd.read_csv(base_path+'input/test/'+event + '-hits.csv')\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 367818,
          "author_name": "returnofsputnik",
          "author_url": "",
          "post_date": "08/08/2018 15:49:24",
          "content": "<p>So you have solved it by reading that data in CSV form, as opposed to making one really big data frame. What if am I restricted to the large dataframe with event ID's in it? I feel that I become stuck because I must somehow only get <code>df.loc[df['event']==i]</code> yet I can't pass <code>df</code>! Hm. Would this make do as a workaround, or would it still make bottlenecks?</p>\n\n<pre><code>def get_hits(i):\n    df_temp = df.loc[df['event']==i]\n    return df_temp\n\n\ndef work_sub(param):\n    (i, &lt;other parameter&gt;) = param\n    hits = get_hits(i)\n</code></pre>\n\n<p>I predict this will still make a bottleneck, however, because to execute <code>get_hits</code> the worker needs to read in <code>df</code>. Therefore I fear the only way to make this work is if your dataset can be read one-by-one, like as if it were stored in CSV format earlier so you can read it in by the event_id. If it's in one large dataframe, it seems I'm screwed . . !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 367822,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/08/2018 15:52:06",
          "content": "<p>Do you get that your way means the big dataframe is copied in each worker?  If you don't then I suggest you study how parallelism works in Python.  If you do then I think you get why the way I propose is better.</p>\n\n<p>Edit: why do you want to load all event data, concatenate it into a gigantic data frame, then have it split back into each worker?  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 367835,
          "author_name": "returnofsputnik",
          "author_url": "",
          "post_date": "08/08/2018 16:15:12",
          "content": "<p>I am learning from your example because I'm implementing parallelism at my work, and I read Kaggle posts for learning. I'm not given each event data as an individual CSV, and I am trying to use your tips to change my current parallelism implementation to make it run faster. I <em>start</em> with one gigantic data frame.</p>\n\n<p>According to this post, <a href=\"https://stackoverflow.com/questions/33612935/large-pandas-dataframe-parallel-processing\">https://stackoverflow.com/questions/33612935/large-pandas-dataframe-parallel-processing</a> it appears that I am forced to let the parallel package pickle my data frame each time a worker wants to access it.</p>\n\n<p>According to this post, <a href=\"https://stackoverflow.com/questions/40357434/pandas-df-iterrow-parallelization\">https://stackoverflow.com/questions/40357434/pandas-df-iterrow-parallelization</a> a partial workaround is to split it up into the number of processes I want to have.</p>\n\n<p>Also recommended was a package called Dask. </p>\n\n<p>For now however, because my dataframe is small (only about 50,000 rows), I think I might change my code to split the dataframe as recommended by the second post.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 367841,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/08/2018 16:30:07",
          "content": "<p>I suggest you first write each event data in a separate file on disk, then you can implement what I describe, loading only the relevant file into each worker.</p>\n\n<p>The first post you point to is also documenting the main issue with joblib: you have to load all data in the master, then it is copied into each worker.  That's why I am not using joblib and rather use Python parallel processing directly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 367873,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/08/2018 18:12:00",
          "content": "<p>Note that what I suggest makes sense for the specific computation we have here, given each event data is in separate files.  If you are discussing parallel processing of data frames in general, then it is different.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 367862,
      "author_name": "jackvial",
      "author_url": "",
      "post_date": "08/08/2018 17:33:59",
      "content": "<p>This seem a bit more elegant than my solution which was to process all events without saving each to disk on a 72 core aws instance with about 144gb of ram. (c5.18xlarge) Although it was able to process the whole submission in 3 hours which was just over twice the time it took to process one event locally.</p>",
      "votes": null,
      "replies": [
        {
          "id": 368221,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "08/09/2018 14:04:50",
          "content": "<p>BTW @Jack Vial, what IDE do you use on cloud? It's a bit out of topic, but related. </p>\n\n<p>I've no IT background at all (I googled how to use github :)), following simple instructions I installed anaconda, but could not start spyder .... then I installed pycharm using umake, but it would not start giving me Pycharm Startup Error: Unable to detect graphics environment. \nI found it: <a href=\"https://stackoverflow.com/questions/46124295/pycharm-startup-error-unable-to-detect-graphics-environment\">https://stackoverflow.com/questions/46124295/pycharm-startup-error-unable-to-detect-graphics-environment</a>\nI tried their solutions.\n1. I tried to install in opt but it tells me I do not have the privilege\n~$ /opt/pycharm/bin/pycharm.sh <br>\n2. Tried this one as well, did not work\n$ chmod +x pycharm.sh , \nrun pycharm with ./pycharm.sh <br>\n3. I install java which is not headless, but the same mistake.  :( I do not know what else to do/try, maybe some other IDE will work out better... \nPS I set up google cloud as it was easier then AWS, chose ubuntu 16.04 as in tutorial. Maybe there is a simple instruction that works for community pycharm, with empty ubunta (maybe I chose a wrong tutorial, i tied two with snap as well, with snap did not work)?   Maybe you could help me a bit with advice as you are the cloud user? I wanted to use cloud for other projects, so need to sort it not only for this competition</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368239,
          "author_name": "jackvial",
          "author_url": "",
          "post_date": "08/09/2018 14:38:45",
          "content": "<p>I write and test my code locally in VSCode then push to github, then ssh into the AWS machine and pull the changes down from github. If I do need to make changes to the code directly on the AWS machine I use the Vim text editor which runs in the terminal  and is usually installed by default on Ubuntu and some other Linux distributions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368242,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "08/09/2018 14:48:46",
          "content": "<p>Yes, i have Vim. I still did not get it, when you take your ready code from github what program do you use to run it? you need some IDE, don't you? (I am not a linux user)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368245,
          "author_name": "jackvial",
          "author_url": "",
          "post_date": "08/09/2018 14:54:18",
          "content": "<p>I run the code from the terminal. Something like <code>python process_submission.py</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368246,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/09/2018 14:55:58",
          "content": "<p>Jack, How did you parallelize code?  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368253,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "08/09/2018 15:11:40",
          "content": "<p>ou, you can do just like that from a terminal, cool... but how it gets all the imports? I put a clusterer class in a separate py file and then do import from it, maybe it's not a good idea... and I use codes from trackml library and import from then load dataset and scoring. How do you handle that? Do you put all that in one single file? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368260,
          "author_name": "jackvial",
          "author_url": "",
          "post_date": "08/09/2018 15:18:49",
          "content": "<p>@CPMP I do something like this:</p>\n\n<pre><code>from trackml.dataset import load_event, load_dataset\nfrom trackml.score import score_event\nfrom joblib import Parallel, delayed\nimport multiprocessing\nimport time\nnum_cores = multiprocessing.cpu_count()\npath_to_test = \"./../data/test\"\n\ndef processOneSubmissionEvent(event_id, hits, cells):\n    model = Clusterer()\n    labels = model.run(hits)\n    submission = create_one_event_submission(event_id, hits, labels)\n    return (event_id, submission)\n\nprocessed_events = Parallel(n_jobs=num_cores)(delayed(processOneSubmissionEvent)(event_id, hits, cells) for event_id, hits, cells in load_dataset(path_to_test, parts=['hits', 'cells']))\n    test_dataset_submissions = []\n    for event in processed_events:\n        test_dataset_submissions.append(event[1])\n\n    # Create submission file\n    ts = str(time.time()).split(\".\")[0]\n    final_submussion = pd.concat(test_dataset_submissions, axis=0)\n    final_submussion.to_csv(\"./../submissions/\" + ts + \"_submission.csv\", index=False)\n    print(\"submission created\")\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368262,
          "author_name": "jackvial",
          "author_url": "",
          "post_date": "08/09/2018 15:20:32",
          "content": "<p><a href=\"/blonde\">@blonde</a> The imports will be imported using the import statements usually at the top of the script e.g.</p>\n\n<p>from trackml.dataset import load_event, load_dataset</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368265,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "08/09/2018 15:28:16",
          "content": "<p>ok, got it, will try. Do I need to put all those progs to a special folder on server where my python exe file is, or it can run from any location? Sorry for this dumb questions, i am totally new to linux, cloud, python, ML... only google and people save me</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368745,
          "author_name": "jackvial",
          "author_url": "",
          "post_date": "08/10/2018 17:04:17",
          "content": "<p><a href=\"/blonde\">@blonde</a> </p>\n\n<blockquote>\n  <p>Do I need to put all those progs to a special folder on server where\n  my python exe file is</p>\n</blockquote>\n\n<p>If you are refereeing to the module imports then python will know where to look for them if they are installed with pip or anaconda. If the module you are importing is some file in your project then it needs to be in the same directory as the file that is importing it or you need to add it to your path so python knows where to look for it.</p>\n\n<blockquote>\n  <p>Sorry for this dumb questions, i am totally new to linux, cloud,\n  python, ML... only google and people save me</p>\n</blockquote>\n\n<p>There are no dumb questions! :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368758,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "08/10/2018 17:46:01",
          "content": "<p>@Jack Vial </p>\n\n<p>I managed to get my first paralleled project in Windows! With all my imports etc. \nWith no IT background I never considered using just cmd, now I see it's better in some ways. Thank you for your code and comments! </p>\n\n<p>The next step is doing it on linux... Maybe it would not help me to improve before the deadline, but I started kaggle to learn things and that was great learning fun! I cannot add emotion here, but you can imagine a little penguin dancing :)))</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368768,
          "author_name": "jackvial",
          "author_url": "",
          "post_date": "08/10/2018 18:27:36",
          "content": "<p><a href=\"/blonde\">@blonde</a> Congrats! Yes the cmd/terminal/shell is very useful and powerful, especially on linux. </p>\n\n<p>Definitely if not in this competition everything you learn will help in the next! Maybe you will need to put in a request for emojis! :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368951,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "08/11/2018 11:29:15",
          "content": "<p>@Jack Vial </p>\n\n<p>I now managed to start it on google cloud!!! 24 cpu, it does not scale like that, it's still two times slower than laptop event time*124/24, but still... I'll ask for the emotions,  but for now you can imagine 24 little penguins dancing :))</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 367884,
      "author_name": "blondinka",
      "author_url": "",
      "post_date": "08/08/2018 18:46:24",
      "content": "<p>thank you for this post! It looks easier then <a href=\"https://pythonhosted.org/joblib/parallel.html\">https://pythonhosted.org/joblib/parallel.html</a> I tried to implement before... we'll see</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 367927,
      "author_name": "sergeyzlobin",
      "author_url": "",
      "post_date": "08/08/2018 21:03:39",
      "content": "<p>I was confused by the parameter 'maxtasksperchild=1'. The documentation says ' A frequent pattern found in other systems (such as Apache, mod_wsgi, etc) to free resources held by workers is to allow a worker within a pool to complete only a set amount of work before being exiting, being cleaned up and a new process spawned to replace the old one.'\nI think it can take to spawn a new process. Why not to use an existing process? Memory fragmetation etc?</p>",
      "votes": null,
      "replies": [
        {
          "id": 368056,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/09/2018 06:43:58",
          "content": "<p>This parameter forces to stop each worker process after it has processed one event, thereby freeing all resources associated with it.  If you don't use it, then you keep <code>n_proc</code> workers resources until the end of computation.  It does not make much difference till the number of remaining events is less than <code>n_proc</code>. At that point the number of workers starts decreasing with this setting, rather than having idle workers.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368339,
          "author_name": "sergeyzlobin",
          "author_url": "",
          "post_date": "08/09/2018 18:16:02",
          "content": "<p>&gt; It does not make much difference till the number of remaining events is less than n_proc.</p>\n\n<p>I use parallelization inside one event, every dbscan is a task. So one task is less than 1 second. I think it will be not efficient to kill processes.\nBy the way, don't you use parallelization in one event? So how do you test weigths etc? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368498,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/10/2018 03:04:05",
          "content": "<p>I use parallelism for events and I don't use parallelism inside each event.  It is a coarse grain parallelism.  You use fine grain parallelism, which makes the master a bottleneck very easily.   My gut feeling is that this is not going to be as efficient as coarse grain parallelism, and it probably explain why it does not scale beyond 3 parallel processes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368765,
          "author_name": "sergeyzlobin",
          "author_url": "",
          "post_date": "08/10/2018 18:05:13",
          "content": "<blockquote>\n  <p>My gut feeling is that this is not going to be as efficient as coarse grain parallelism, and it probably explain why it does not scale beyond 3 parallel processes.</p>\n</blockquote>\n\n<p>Yes, probably you're right. Anyway I don't have machines with more than 4 cores. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368969,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/11/2018 12:47:35",
          "content": "<blockquote>\n  <p>I don't have machines with more than 4 cores. </p>\n</blockquote>\n\n<p>Yes, but you're using only 3.  If you could use 4 then you would get a 33% speedup ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 367936,
      "author_name": "jaybob20",
      "author_url": "",
      "post_date": "08/08/2018 21:31:05",
      "content": "<p>This is what I use for all the version of Clusterer() out there. I doesn't seem to use much memory. I guess if you had more cores you would have more memory too?</p>\n\n<pre><code>def worker(event_id, hits, cells):\n    model = Clusterer()\n    labels = model.predict(hits)\n    print('Event ID: ', event_id)\n    # Prepare submission for an event\n    one_submission = create_one_event_submission(event_id, hits, labels)\n    return one_submission\n\nfrom multiprocessing import Pool\n\npool = Pool(processes=6)\ntest_dataset_submissions = pool.starmap(worker,load_dataset(path_to_test, parts=['hits', 'cells']))\n\n# Create submission file\nsubmission = pd.concat(test_dataset_submissions, axis=0)\nsubmission.to_csv('submission.csv', index=False)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 367951,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "08/08/2018 23:26:00",
          "content": "<p>I am new to python, sorry for this stupid questions: \n1. I guess @CPMP suggests to place loader in the worker, it speeds up, and he is using his hits = pd.read_csv() instead of trackml libruary </p>\n\n<ol>\n<li>pool.starmap(worker,load_dataset(path_to_test, parts=['hits', 'cells'])) \nshould it be worker(event_id, hits, cells = load_dataset(path_to_test, parts=['hits', 'cells'])) </li>\n</ol>\n\n<p>should it be the for loop for events from 1 to 125 somewhere? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 367960,
          "author_name": "jaybob20",
          "author_url": "",
          "post_date": "08/09/2018 00:27:28",
          "content": "<p>The load time isn't my bottleneck, so I'm loading from the main part, and if all the process start loading files at once it could create another slowdown... you can never win with parallelism.\nload_dataset(path_to_test, parts=['hits', 'cells']) gives you a list of tuples like this:</p>\n\n<pre><code>[(0,hits,cells),(1,hits,cells),...]\n</code></pre>\n\n<p>so if you just pass the tuple to the worker function you would have to handle that some way. Or you pass with starmap and it sends 3 things to the worker instead.\nHope that clears it up.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368057,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/09/2018 06:50:50",
          "content": "<p>&gt; The load time isn't my bottleneck</p>\n\n<p>How do you know that?</p>\n\n<p>Anyway, your way of doing thing is the one that is slowest as I explained, because you force all data to be loaded by the master, then copied to each worker.  It does not matter much if you use a small number of workers, but will degrade computation time if you use many.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368163,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "08/09/2018 11:46:27",
          "content": "<p>@Mads I am still on the way with parallelism (new thing for me totally)... Looking at CPMP code i added maxtasksperchild=1 to yours to make sure each worker process one task. I got my spyder hanged.... and I added lines like print(\"start worker\"), so it hanged before any progress lines, which make me think it's loading. But maybe i made a mistake somewhere else.\nso how long load_dataset time is approx for you? I try to implement loading in the worker in the meantime</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368187,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/09/2018 12:48:26",
          "content": "<p>@Blonde it hangs because the master is busy loading data for the workers.  Or you are working on Windows and inside notebooks, in which case you are in trouble :(</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368237,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "08/09/2018 14:33:54",
          "content": "<p>@CPMP I implemented your code and it also hangs... I googled and found it is a spyder problem: <a href=\"https://github.com/spyder-ide/spyder/issues/2937\">https://github.com/spyder-ide/spyder/issues/2937</a>. \nWhat IDE do you use for successful parallel computing in python? Will jupyter work? Or better community pycharm?  Or better something else? \nI removed all other parameters, so it's not the issue. </p>\n\n<p>Also, this line is not a pseudo code, it's a valid line, correct? params = [i for i in range(10)]  I just never met such syntax before)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368240,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "08/09/2018 14:39:26",
          "content": "<p>@CPMP yes, I am on Windows and inside spyder... I still cannot set up pycharm on ubuntu 16, used few tutorials and get Pycharm Startup Error: Unable to detect graphics environment, reinstalled, reinstalled java, googled that problem and tried all solutions I found, and still no progress...  I should rename my account to Dumb Blonde :))). What do you use for python on linux, is it ubunta? Can I do parallel computing in windows 7 or it's not gonna work? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368244,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/09/2018 14:54:12",
          "content": "<p>I never managed to get Python parallelism work on Windows, but I was not as fluent with Python as I am now.</p>\n\n<p>On ubuntu why don't you use Jupyter Notebook?  It is quite convenient for data science and machine learning.  I use ubuntu 14.04, or ubuntu 16.04 on Intel machines, and RHEL on a Power machine.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368252,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "08/09/2018 15:08:11",
          "content": "<p>@CPMP Thank you. I installed Jupyter Notebook on google cloud to learn fastai course. It starts a kernel but then kernel dead. \"The kernel has died, and the automatic restart has failed\". Found the story with the mistakes, troubleshooting... I think you can add to your discussion for newbies -- USE LINUX !!! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368261,
          "author_name": "jaybob20",
          "author_url": "",
          "post_date": "08/09/2018 15:19:17",
          "content": "<p>@CPMP Load time from disk for all the data is 30 seconds for me. 27 seconds after it is cached (I'm on Ubuntu). Lets say the data is loaded and most of the 27 seconds is making dataframes. It has to pass these dataframes to the workers, so we have another memory copy. Lets just say it takes the same amount of time and double the data time. That is still less than 1 sec per event. My model runs a little slower than 1 sec per event.</p>\n\n<p>@Blonde You can see my load time above. Also load_dataset is a generator, it only loads or does something when you try to get something from it.\nif you do something like:</p>\n\n<pre><code>a=load_dataset(path_to_test, parts=['hits', 'cells'])\n</code></pre>\n\n<p>nothing happens, it just gets thing ready to use.\nI really see no need in setting maxtasksperchild. It could make things slower if you had the same data that was used in each child. I would use it if the event data was dramatically different in size per event and memory was an issue, sometimes the actual python task doesn't like to let go of ram or some lib has a memory leak.</p>\n\n<p>I've ran this code successfully in notebook and as a script. I'm running 10 threads on 6 core i7, there is about a 30 second pause before the works start churning and they all seem to start working at once. So my calculations above could be off by a second and data load is take 3 seconds per event. With 10 threads I'm using less than 2GB of ram.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368284,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "08/09/2018 16:09:12",
          "content": "<p>Thank you @Mads. I use your pool.starmap(worker,load_dataset(path_to_test, parts=['hits', 'cells'])), but i save submission for each event individually in worker. My problem is Windows 7. Now struggling to set up jupyter on linux cloud... installed but kernel dead :(, doing troubleshooting, will see ... Now I see that kaggle is for linux users :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368341,
          "author_name": "jaybob20",
          "author_url": "",
          "post_date": "08/09/2018 18:21:38",
          "content": "<p>Check out Docker running a linux/data science container on windows. It might be easier. I haven't set up  personally for this, but on other things on a mac.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368395,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/09/2018 20:11:32",
          "content": "<blockquote>\n  <p>My model runs a little slower than 1 sec per event.</p>\n</blockquote>\n\n<p>@Mads,  you don't seem to have a clear view of how parallel computing works here.  First, you are off by a factor of 40 or so on the overhead for data loading.  indeed, here is how things run:</p>\n\n<ol>\n<li><p>The master loads all data.  This takes 27 seconds.</p></li>\n<li><p>The master dispatches the data to workers.  It does it one worker at a time because of the GIL.  You say this takes about the same time, which means that on average workers get their data in 27/2 seconds.  As a result, workers wait 3/2 * 27 seconds on average, i.e. 40 seconds, not one second.</p></li>\n</ol>\n\n<p>Granted, this is not a big deal if workers runs for hours.</p>\n\n<p>But if workers run for hours then you have another issue.  With default settings, the Pool assigns more than one tasks per worker.  With 125 tasks and 20 workers (my case), it assigns 2 tasks per worker.  This means that after processing the first 120 events, it assigns the last 5 events to 3 workers, not 5.  It means that processing the last 5 events takes twice the time to process one event, even if you have more than 5 workers.  It my case it means that the overall computation takes about 8 times the time it takes to process one event, instead of 7. it means 10 extra hours.  Not negligible at all.</p>\n\n<p>This said, if you think your way is fine, great.  Do as you wish.  But you posted your (bad) way of doing things, which is why I put the dots on the i.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368417,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "08/09/2018 21:19:09",
          "content": "<p>ok, while I unsuccessfully try to install everything on google cloud i drove to Uni where they have linux. Now both variants start, but for @Mads version i have:\n'Pool' object has no attribute 'starmap'\nand for @CPMP version I have: 'total' is an invalid keyword argument for this function (for ls   = pool.map(work_sub, params, chunksize=1) )\nI think it's because there is python 2.7 there and maybe older multiprocessing and Pool... but i cannot update it as it's not my server... I keep optimizing my clustering, hopefully I'll be able to sort it out and run a final submission </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368419,
          "author_name": "jaybob20",
          "author_url": "",
          "post_date": "08/09/2018 21:23:23",
          "content": "<p>@CPMP If I have time I will do the load in the worker and time each run. Like I said at the end post it takes about 30 seconds for things to get going. After that I don't see any system load decrease until things taper off at the end. And it is clear each process is running only one event.</p>\n\n<blockquote>\n  <p>maxtasksperchild is the number of tasks a worker process can complete before it will exit and be replaced with a fresh worker process, to enable unused resources to be freed. The default maxtasksperchild is None, which means worker processes will live as long as the pool.</p>\n</blockquote>\n\n<p>This has nothing to do with \"chunksize.\" I think you might be confusing the two.</p>\n\n<blockquote>\n  <p>map(func, iterable[, chunksize])\n  A parallel equivalent of the map() built-in function (it supports only one iterable argument though). It blocks until the result is ready.</p>\n  \n  <p>This method chops the iterable into a number of chunks which it submits to the process pool as separate tasks. The (approximate) size of these chunks can be specified by setting chunksize to a positive integer.</p>\n</blockquote>\n\n<p>And your math doesn't seem to match my results. If it was 40 seconds per task, it would take 400 seconds for things to start working for 10 tasks, it takes 30 seconds as I mentioned. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368420,
          "author_name": "jaybob20",
          "author_url": "",
          "post_date": "08/09/2018 21:27:43",
          "content": "<p>@Blonde starmap is in python3.3 and on. You could see if python3 is on the system by running your script with python3 your_script.py\nIf that doesn't do it there is always setting up a virtual_env and having python3 in that.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368497,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/10/2018 03:01:04",
          "content": "<blockquote>\n  <p>If it was 40 seconds per task, it would take 400 seconds for things to start working for 10 tasks,</p>\n</blockquote>\n\n<p>You should read more carefully.  I wrote that each of your worker would start roughly 40 second after the master starts.  </p>\n\n<p>Anyway, \"on ne fait pas boir un ane qui n'a pas soif\".  Sorry for the French, I don't know the English equivalent.  </p>\n\n<p>As I said, if you're fine with your code, fine with me.  I know for sure, because I measured it, that my way saves me 10 hours per submission. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368731,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "08/10/2018 16:23:00",
          "content": "<p>@CPMP I managed to parallel in Windows! It works in cmd, the trick is you need to add (otherwise, it hangs): \nif <strong>name</strong> == '<strong>main</strong>':\n    if use_parallel: \n        print('start worker')\n        pool = Pool(processes=n_proc, maxtasksperchild=1)\n        ls   = pool.map( work_sub, params, chunksize=1 )\n        pool.close()\n    else:\n        ls = [work_sub(param) for param in params]</p>\n\n<p>I never considered using just cmd, now I see it's better in some ways. Thank you for your code, and thanks to @Mads and  @Jack Vial for comments. Next step is doing it on linux... Maybe it would not help me to improve before the deadline, but I started kaggle to learn things and it was great learning fun! I cannot add emotion here, but you can imagine a little penguin dancing :))) </p>\n\n<p>PS. @CPMP I really appreciate this post! The only little thing: such obvious things for you and other experienced users are nice to publish a bit earlier, like a month before the deadline, for the complete beginners, who take time to make it and did not hear about parallel programming</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368762,
          "author_name": "sergeyzlobin",
          "author_url": "",
          "post_date": "08/10/2018 17:57:57",
          "content": "<p>&gt; I managed to parallel in Windows! It works in cmd, the trick is you need to add (otherwise, it hangs): if name == 'main': </p>\n\n<p>Yeah, you can look at my kernel with parallelization (not events, but dbscans). I use Windows too.</p>\n\n<p><a href=\"https://www.kaggle.com/sergeyzlobin/unrolling-helices-baseline-python\">https://www.kaggle.com/sergeyzlobin/unrolling-helices-baseline-python</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368771,
          "author_name": "shivrajp",
          "author_url": "",
          "post_date": "08/10/2018 18:55:51",
          "content": "<p>I just prefer to use windows subsystem for linux. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368801,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/10/2018 21:04:54",
          "content": "<blockquote>\n  <p>PS. @CPMP I really appreciate this post! The only little thing: such obvious things for you and other experienced users are nice to publish a bit earlier, like a month before the deadline, for the complete beginners, who take time to make it and did not hear about parallel programming</p>\n</blockquote>\n\n<p>Better late than never ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368949,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "08/11/2018 11:27:52",
          "content": "<p>@CPMP Better late than never ;)  True...  I now managed to start it on google cloud!!! 24 cpu (I was stupid and chose a wrong region instead of usa), so now running on linux...it should help me better. I thought it would be 24 times faster than laptop (yes, I did not prallel my calculations before, at all... ), it does not scale like that, it's still two times slower than laptop event time*124/24, but still...  Can't add emotion, but now you can imagine 24 little penguins dancing :))</p>\n\n<p>@Sergey, I saw your kernel a while ago but was unable to understand it (I've zero IT background), so now as I got my better submissions running I was planning exactly that: parallel in the main to optimise things better... so much more I could have done if I knew about parallel... optimising things much much faster and for many different features, and on 10 events, instead of picking one, waiting for 30 min for a single calculation and hoping for the best :)... Anyway, with my zero IT/linux background bronze is ok for the debut.  And now I feel I'll be able to stay there.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 368337,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "08/09/2018 18:14:02",
      "content": "<p>Our current solution (0.646) runs in R and takes about 150 minutes to create the testset submission using lots of multiprocessing techniques :-)</p>",
      "votes": null,
      "replies": [
        {
          "id": 368360,
          "author_name": "sergeyzlobin",
          "author_url": "",
          "post_date": "08/09/2018 19:04:50",
          "content": "<p>R is better (maybe for multiprocessing  )? Why not python? </p>\n\n<p>It takes 2 or 3 days for me to create the testset submission. :((</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368376,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "08/09/2018 19:36:10",
          "content": "<p>Sergey, same for me 3 days for the miserable score on a usual windows laptop (cause I did not know about linux and that I can parallel...). I still struggle to set python IDE on linux, but I keep trying... now uninstalled everything reinstalled... and of course, nothing works :) This ML problem turned our to be linux problem for me :))) but then i can enjoy running progs on linux afterwards (if i set it up)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368422,
          "author_name": "titericz",
          "author_url": "",
          "post_date": "08/09/2018 21:33:55",
          "content": "<p>R is not better than python for multiprocessing.\nI just figured out that my algorithm, for some reason that I don't know, runs faster in R :-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368494,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/10/2018 02:57:45",
          "content": "<blockquote>\n  <p>runs faster in R :-)</p>\n</blockquote>\n\n<p>Wow! ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 369118,
      "author_name": "bkkaggle",
      "author_url": "",
      "post_date": "08/12/2018 02:49:41",
      "content": "<p>When using multiple processes, Python copies global variables to each worker and doesn't sync any changes made to these global variables back to the master process since all the processes are concurrent. Each process could merge the current track and a global variable containing the best track, but each process wouldn't be able to update the global variable containing the best track. How do you make it work?</p>",
      "votes": null,
      "replies": [
        {
          "id": 369153,
          "author_name": "shivrajp",
          "author_url": "",
          "post_date": "08/12/2018 06:16:11",
          "content": "<p>You can use Manger list from multiprocessing. This should help you manage global variables across multi processes.</p>\n\n<p><code>from multiprocessing import Manager</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 369186,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/12/2018 08:51:28",
          "content": "<p>You can use the <code>work_sub</code> return value for passing results back to the master.  The list <code>ls</code> will contain all the returned values.  Or you can have the master code read the saved event submissions.  This is what I do after the parallel loop is executed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 369239,
          "author_name": "bkkaggle",
          "author_url": "",
          "post_date": "08/12/2018 13:39:46",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> </p>\n\n<blockquote>\n  <p>you can have the master code read the saved event submissions</p>\n</blockquote>\n\n<p>This is what I currently do; Whats different is that I parallelize over each DBSCAN iteration while the code above parallelizes over each event. I think that if I'm using this approach, using multiprocessing.manager should solve the problem. I guess that if you're running each event on a different event, your DBSCAN n_jobs value is 1. Would it be a good idea to use part of the cores to parallelize over the events and part of the cores to parallelize over the DBSCAN iterations?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 369245,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/12/2018 13:49:06",
          "content": "<p>TL;DR Do you see a linear speedup as you increase the number of workers in you case?  I see linear speedup for 40 cores.</p>\n\n<p>Parallelism has some overhead, and the coarser the task, the better.  It is why I use parallelism per event and computation for one event is single threaded.  This way scales almost perfectly for 40 cores.   I don't see why one would need more parallelism unless one has more cores than events...</p>\n\n<p>Finer grained parallelism is better handle via multi threading, but Python does not support multi threading.  Sure, there is a package for it, but in general only one thread can execute at a time because of the GIL.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 369292,
          "author_name": "johnhsweeney",
          "author_url": "",
          "post_date": "08/12/2018 16:19:52",
          "content": "<p>I used \"preemptable\" server instances from Google Cloud (to save money...).  Since you can expect to be preempted a couple of times per day, I chose to use parallelism to finish one event as quickly as possible, even if it isn't the most efficient method.  Writing code to recover one event from an interruption seemed easier than code to recover 90+ ... but maybe I was just being lazy!</p>\n\n<p>My local code was written to process one event per thread, but as my solution complexity grew where I couldn't generate a submission in a couple of days on a 10-core CPU, I moved to the cloud.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 369295,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/12/2018 16:28:12",
          "content": "<p>John, this is indeed a good reason to use parallelism within one event computation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "367749": "As many noticed, it is possible to speed up submission computation using parallel processing, with a master process and a number of worker processes.  Python supports it, but there are few tricks to make it efficient.\n\n1. Do not pass data to the workers, rather have each worker load its data from disk.  Indeed, passing data means the master process must load it first, and this can become a real bottleneck for the whole computation.\n\n2. Make sure each worker works on a single event.  Indeed, Python Pool assigns more than one task to each worker with its default settings. This can result in some of your processors becoming idle while events remain to be processed.\n\nCode to do it in Python 3.6 is given below.  It assumes you implement the computation of one event submission  as a single function `work_sub` that gets one parameter, namely the event number.  You can pass additional parameters but be careful, all data you pass there can create a bottleneck for the whole computation, stick to simple parameters like numerical values.  Do not pass the event data.\n\nThen in pseudo code:\n\n    from multiprocessing import Pool\n    \n    def work_sub(param):\n        (i,",
    "367780": "&gt; passing data means the master process must load it first, and this can become a real bottleneck for the whole computation.\n\nDo you know a time for passing data? Suppose 1 dbscan is about 1 second. Is it comparable to it? I pass 'hits' values and other params. I thought passing is fast. :(",
    "367783": "The problem is not that passing is fast or not.  The problem is that python cannot execute more than one thread at a time because of the GIL.  It means that the master can pass data to only one worker at a time, making it non parallel...  The effect is worse when your worker runs fast, or if you have many workers.",
    "367790": "The worst I can imagine is to load all events data in the master, then pass it to each worker.  Then the worker code selects the relevant part of it.  \n\nWith this pattern you load the 125 events data `n_proc + 1` times, overloading your machine memory, and slowing it dramatically.",
    "367800": "\"You can pass additional parameters but be careful, all data you pass there can create a bottleneck for the whole computation, stick to simple parameters like numerical values. \" - Ok, but then if you are required to get the data for event i, you are still required to access the event data by doing something like `df.loc[df['event']==i]`. So you are forced to pass in `df`, unless you make `df` a global variable",
    "367811": "You are describing one of the worst, if not the worst, way of implementing parallelism: you make all event data copied in each worker.\n\n&gt; you are still required to access the event data by doing something like df.loc[df['event']==i]\n\nAbsolutely not, you can load each event data independently in each worker.  For instance, here is the start of my worker code:\n\n    def get_event(i):\n        prefix = 'event000000'\n        if i &lt; 100:\n            prefix = prefix +'0'\n        if i &lt; 10:\n            prefix = prefix +'0'\n        return prefix + str(i)\n    \n    base_path = '/home/jfpuget/Kaggle/TrackML/'\n    \n    def work_sub(param):\n        (i,",
    "367818": "So you have solved it by reading that data in CSV form, as opposed to making one really big data frame. What if am I restricted to the large dataframe with event ID's in it? I feel that I become stuck because I must somehow only get `df.loc[df['event']==i]` yet I can't pass `df`! Hm. Would this make do as a workaround, or would it still make bottlenecks?\n\n    def get_hits(i):\n        df_temp = df.loc[df['event']==i]\n        return df_temp\n\n\n    def work_sub(param):\n        (i,",
    "367821": "I've taken measurements in my case. The overhead of parallelization is about 8% (including parameter passing). A lot, but not deadly.\n\nAnd thanks for this info! I didn't think about this overhead.",
    "367822": "Do you get that your way means the big dataframe is copied in each worker?  If you don't then I suggest you study how parallelism works in Python.  If you do then I think you get why the way I propose is better.\n\nEdit: why do you want to load all event data, concatenate it into a gigantic data frame, then have it split back into each worker?",
    "367835": "I am learning from your example because I'm implementing parallelism at my work, and I read Kaggle posts for learning. I'm not given each event data as an individual CSV, and I am trying to use your tips to change my current parallelism implementation to make it run faster. I *start* with one gigantic data frame.\n\nAccording to this post, https://stackoverflow.com/questions/33612935/large-pandas-dataframe-parallel-processing it appears that I am forced to let the parallel package pickle my data frame each time a worker wants to access it.\n\nAccording to this post, https://stackoverflow.com/questions/40357434/pandas-df-iterrow-parallelization a partial workaround is to split it up into the number of processes I want to have.\n\nAlso recommended was a package called Dask. \n\nFor now however, because my dataframe is small (only about 50,000 rows), I think I might change my code to split the dataframe as recommended by the second post.",
    "367841": "I suggest you first write each event data in a separate file on disk, then you can implement what I describe, loading only the relevant file into each worker.\n\nThe first post you point to is also documenting the main issue with joblib: you have to load all data in the master, then it is copied into each worker.  That's why I am not using joblib and rather use Python parallel processing directly.",
    "367843": "Parallelism overhead is so small it is not measurable with my code ;)  I agree that 8% is not deadly, I guess you don't use many parallel processes.",
    "367862": "This seem a bit more elegant than my solution which was to process all events without saving each to disk on a 72 core aws instance with about 144gb of ram. (c5.18xlarge) Although it was able to process the whole submission in 3 hours which was just over twice the time it took to process one event locally.",
    "367873": "Note that what I suggest makes sense for the specific computation we have here, given each event data is in separate files.  If you are discussing parallel processing of data frames in general, then it is different.",
    "367874": "&gt; I guess you don't use many parallel processes. \n\nYou're right. I use 3 processes (sometimes on different machines). If use all 4 cores, all becomes slow and I can't do anything else.",
    "367876": "See, your way is not enabling you to use your 4 cores...  I bet you would be able to use them if you move to the way I describe above.\n\nYour overhead is way higher than you think, you should measure it when running with 4 workers (one per core).",
    "367884": "thank you for this post! It looks easier then https://pythonhosted.org/joblib/parallel.html I tried to implement before... we'll see",
    "367927": "I was confused by the parameter 'maxtasksperchild=1'. The documentation says ' A frequent pattern found in other systems (such as Apache, mod_wsgi, etc) to free resources held by workers is to allow a worker within a pool to complete only a set amount of work before being exiting, being cleaned up and a new process spawned to replace the old one.'\nI think it can take to spawn a new process. Why not to use an existing process? Memory fragmetation etc?",
    "367936": "This is what I use for all the version of Clusterer() out there. I doesn't seem to use much memory. I guess if you had more cores you would have more memory too?\n\n\n    def worker(event_id, hits, cells):\n        model = Clusterer()\n        labels = model.predict(hits)\n        print('Event ID: ', event_id)\n        # Prepare submission for an event\n        one_submission = create_one_event_submission(event_id, hits, labels)\n        return one_submission\n    \n    from multiprocessing import Pool\n    \n    pool = Pool(processes=6)\n    test_dataset_submissions = pool.starmap(worker,load_dataset(path_to_test, parts=['hits', 'cells']))\n    \n    # Create submission file\n    submission = pd.concat(test_dataset_submissions, axis=0)\n    submission.to_csv('submission.csv', index=False)",
    "367951": "I am new to python, sorry for this stupid questions: \n1. I guess @CPMP suggests to place loader in the worker, it speeds up, and he is using his hits = pd.read_csv() instead of trackml libruary \n\n2. pool.starmap(worker,load_dataset(path_to_test, parts=['hits', 'cells'])) \nshould it be worker(event_id, hits, cells = load_dataset(path_to_test, parts=['hits', 'cells'])) \n\nshould it be the for loop for events from 1 to 125 somewhere?",
    "367960": "The load time isn't my bottleneck, so I'm loading from the main part, and if all the process start loading files at once it could create another slowdown... you can never win with parallelism.\nload_dataset(path_to_test, parts=['hits', 'cells']) gives you a list of tuples like this:\n\n    [(0,hits,cells),(1,hits,cells),...]\n\nso if you just pass the tuple to the worker function you would have to handle that some way. Or you pass with starmap and it sends 3 things to the worker instead.\nHope that clears it up.",
    "368056": "This parameter forces to stop each worker process after it has processed one event, thereby freeing all resources associated with it.  If you don't use it, then you keep `n_proc` workers resources until the end of computation.  It does not make much difference till the number of remaining events is less than `n_proc`. At that point the number of workers starts decreasing with this setting, rather than having idle workers.",
    "368057": "&gt; The load time isn't my bottleneck\n\nHow do you know that?\n\nAnyway, your way of doing thing is the one that is slowest as I explained, because you force all data to be loaded by the master, then copied to each worker.  It does not matter much if you use a small number of workers, but will degrade computation time if you use many.",
    "368163": "Mads I am still on the way with parallelism (new thing for me totally)... Looking at CPMP code i added maxtasksperchild=1 to yours to make sure each worker process one task. I got my spyder hanged.... and I added lines like print(\"start worker\"), so it hanged before any progress lines, which make me think it's loading. But maybe i made a mistake somewhere else.\nso how long load_dataset time is approx for you? I try to implement loading in the worker in the meantime",
    "368187": "Blonde it hangs because the master is busy loading data for the workers.  Or you are working on Windows and inside notebooks, in which case you are in trouble :(",
    "368221": "BTW @Jack Vial, what IDE do you use on cloud? It's a bit out of topic, but related. \n\nI've no IT background at all (I googled how to use github :)), following simple instructions I installed anaconda, but could not start spyder .... then I installed pycharm using umake, but it would not start giving me Pycharm Startup Error: Unable to detect graphics environment. \nI found it: https://stackoverflow.com/questions/46124295/pycharm-startup-error-unable-to-detect-graphics-environment\nI tried their solutions.\n1. I tried to install in opt but it tells me I do not have the privilege\n~$ /opt/pycharm/bin/pycharm.sh  \n2. Tried this one as well, did not work\n$ chmod +x pycharm.sh , \nrun pycharm with ./pycharm.sh   \n3. I install java which is not headless, but the same mistake.  :( I do not know what else to do/try, maybe some other IDE will work out better... \nPS I set up google cloud as it was easier then AWS, chose ubuntu 16.04 as in tutorial. Maybe there is a simple instruction that works for community pycharm, with empty ubunta (maybe I chose a wrong tutorial, i tied two with snap as well, with snap did not work)?   Maybe you could help me a bit with advice as you are the cloud user? I wanted to use cloud for other projects, so need to sort it not only for this competition",
    "368237": "CPMP I implemented your code and it also hangs... I googled and found it is a spyder problem: https://github.com/spyder-ide/spyder/issues/2937. \nWhat IDE do you use for successful parallel computing in python? Will jupyter work? Or better community pycharm?  Or better something else? \nI removed all other parameters, so it's not the issue. \n\nAlso, this line is not a pseudo code, it's a valid line, correct? params = [i for i in range(10)]  I just never met such syntax before)",
    "368239": "I write and test my code locally in VSCode then push to github, then ssh into the AWS machine and pull the changes down from github. If I do need to make changes to the code directly on the AWS machine I use the Vim text editor which runs in the terminal  and is usually installed by default on Ubuntu and some other Linux distributions.",
    "368240": "CPMP yes, I am on Windows and inside spyder... I still cannot set up pycharm on ubuntu 16, used few tutorials and get Pycharm Startup Error: Unable to detect graphics environment, reinstalled, reinstalled java, googled that problem and tried all solutions I found, and still no progress...  I should rename my account to Dumb Blonde :))). What do you use for python on linux, is it ubunta? Can I do parallel computing in windows 7 or it's not gonna work?",
    "368242": "Yes, i have Vim. I still did not get it, when you take your ready code from github what program do you use to run it? you need some IDE, don't you? (I am not a linux user)",
    "368244": "I never managed to get Python parallelism work on Windows, but I was not as fluent with Python as I am now.\n\nOn ubuntu why don't you use Jupyter Notebook?  It is quite convenient for data science and machine learning.  I use ubuntu 14.04, or ubuntu 16.04 on Intel machines, and RHEL on a Power machine.",
    "368245": "I run the code from the terminal. Something like `python process_submission.py`",
    "368246": "Jack, How did you parallelize code?",
    "368252": "CPMP Thank you. I installed Jupyter Notebook on google cloud to learn fastai course. It starts a kernel but then kernel dead. \"The kernel has died, and the automatic restart has failed\". Found the story with the mistakes, troubleshooting... I think you can add to your discussion for newbies -- USE LINUX !!!",
    "368253": "ou, you can do just like that from a terminal, cool... but how it gets all the imports? I put a clusterer class in a separate py file and then do import from it, maybe it's not a good idea... and I use codes from trackml library and import from then load dataset and scoring. How do you handle that? Do you put all that in one single file?",
    "368260": "CPMP I do something like this:\n\n    from trackml.dataset import load_event, load_dataset\n    from trackml.score import score_event\n    from joblib import Parallel, delayed\n    import multiprocessing\n    import time\n    num_cores = multiprocessing.cpu_count()\n    path_to_test = \"./../data/test\"\n    \n    def processOneSubmissionEvent(event_id, hits, cells):\n        model = Clusterer()\n        labels = model.run(hits)\n        submission = create_one_event_submission(event_id, hits, labels)\n        return (event_id, submission)\n    \n    processed_events = Parallel(n_jobs=num_cores)(delayed(processOneSubmissionEvent)(event_id, hits, cells) for event_id, hits, cells in load_dataset(path_to_test, parts=['hits', 'cells']))\n        test_dataset_submissions = []\n        for event in processed_events:\n            test_dataset_submissions.append(event[1])\n            \n        # Create submission file\n        ts = str(time.time()).split(\".\")[0]\n        final_submussion = pd.concat(test_dataset_submissions, axis=0)\n        final_submussion.to_csv(\"./../submissions/\" + ts + \"_submission.csv\", index=False)\n        print(\"submission created\")",
    "368261": "CPMP Load time from disk for all the data is 30 seconds for me. 27 seconds after it is cached (I'm on Ubuntu). Lets say the data is loaded and most of the 27 seconds is making dataframes. It has to pass these dataframes to the workers, so we have another memory copy. Lets just say it takes the same amount of time and double the data time. That is still less than 1 sec per event. My model runs a little slower than 1 sec per event.\n\n@Blonde You can see my load time above. Also load_dataset is a generator, it only loads or does something when you try to get something from it.\nif you do something like:\n\n    a=load_dataset(path_to_test, parts=['hits', 'cells'])\n\nnothing happens, it just gets thing ready to use.\nI really see no need in setting maxtasksperchild. It could make things slower if you had the same data that was used in each child. I would use it if the event data was dramatically different in size per event and memory was an issue, sometimes the actual python task doesn't like to let go of ram or some lib has a memory leak.\n \nI've ran this code successfully in notebook and as a script. I'm running 10 threads on 6 core i7, there is about a 30 second pause before the works start churning and they all seem to start working at once. So my calculations above could be off by a second and data load is take 3 seconds per event. With 10 threads I'm using less than 2GB of ram.",
    "368262": "blonde The imports will be imported using the import statements usually at the top of the script e.g.\n\nfrom trackml.dataset import load_event, load_dataset",
    "368265": "ok, got it, will try. Do I need to put all those progs to a special folder on server where my python exe file is, or it can run from any location? Sorry for this dumb questions, i am totally new to linux, cloud, python, ML... only google and people save me",
    "368284": "Thank you @Mads. I use your pool.starmap(worker,load_dataset(path_to_test, parts=['hits', 'cells'])), but i save submission for each event individually in worker. My problem is Windows 7. Now struggling to set up jupyter on linux cloud... installed but kernel dead :(, doing troubleshooting, will see ... Now I see that kaggle is for linux users :)",
    "368337": "Our current solution (0.646) runs in R and takes about 150 minutes to create the testset submission using lots of multiprocessing techniques :-)",
    "368339": "&gt; It does not make much difference till the number of remaining events is less than n_proc.\n\nI use parallelization inside one event, every dbscan is a task. So one task is less than 1 second. I think it will be not efficient to kill processes.\nBy the way, don't you use parallelization in one event? So how do you test weigths etc?",
    "368341": "Check out Docker running a linux/data science container on windows. It might be easier. I haven't set up  personally for this, but on other things on a mac.",
    "368360": "R is better (maybe for multiprocessing  )? Why not python? \n\nIt takes 2 or 3 days for me to create the testset submission. :((",
    "368376": "Sergey, same for me 3 days for the miserable score on a usual windows laptop (cause I did not know about linux and that I can parallel...). I still struggle to set python IDE on linux, but I keep trying... now uninstalled everything reinstalled... and of course, nothing works :) This ML problem turned our to be linux problem for me :))) but then i can enjoy running progs on linux afterwards (if i set it up)",
    "368395": "&gt; My model runs a little slower than 1 sec per event.\n\n@Mads,  you don't seem to have a clear view of how parallel computing works here.  First, you are off by a factor of 40 or so on the overhead for data loading.  indeed, here is how things run:\n\n 1. The master loads all data.  This takes 27 seconds.\n\n 2.  The master dispatches the data to workers.  It does it one worker at a time because of the GIL.  You say this takes about the same time, which means that on average workers get their data in 27/2 seconds.  As a result, workers wait 3/2 * 27 seconds on average, i.e. 40 seconds, not one second.\n\nGranted, this is not a big deal if workers runs for hours.\n\nBut if workers run for hours then you have another issue.  With default settings, the Pool assigns more than one tasks per worker.  With 125 tasks and 20 workers (my case), it assigns 2 tasks per worker.  This means that after processing the first 120 events, it assigns the last 5 events to 3 workers, not 5.  It means that processing the last 5 events takes twice the time to process one event, even if you have more than 5 workers.  It my case it means that the overall computation takes about 8 times the time it takes to process one event, instead of 7. it means 10 extra hours.  Not negligible at all.\n\nThis said, if you think your way is fine, great.  Do as you wish.  But you posted your (bad) way of doing things, which is why I put the dots on the i.",
    "368417": "ok, while I unsuccessfully try to install everything on google cloud i drove to Uni where they have linux. Now both variants start, but for @Mads version i have:\n'Pool' object has no attribute 'starmap'\nand for @CPMP version I have: 'total' is an invalid keyword argument for this function (for ls   = pool.map(work_sub, params, chunksize=1) )\nI think it's because there is python 2.7 there and maybe older multiprocessing and Pool... but i cannot update it as it's not my server... I keep optimizing my clustering, hopefully I'll be able to sort it out and run a final submission",
    "368419": "CPMP If I have time I will do the load in the worker and time each run. Like I said at the end post it takes about 30 seconds for things to get going. After that I don't see any system load decrease until things taper off at the end. And it is clear each process is running only one event.\n\n &gt;maxtasksperchild is the number of tasks a worker process can complete before it will exit and be replaced with a fresh worker process, to enable unused resources to be freed. The default maxtasksperchild is None, which means worker processes will live as long as the pool.\n\nThis has nothing to do with \"chunksize.\" I think you might be confusing the two.\n\n&gt;map(func, iterable[, chunksize])\n&gt;A parallel equivalent of the map() built-in function (it supports only one iterable argument though). It blocks until the result is ready.\n\n&gt;This method chops the iterable into a number of chunks which it submits to the process pool as separate tasks. The (approximate) size of these chunks can be specified by setting chunksize to a positive integer.\n\nAnd your math doesn't seem to match my results. If it was 40 seconds per task, it would take 400 seconds for things to start working for 10 tasks, it takes 30 seconds as I mentioned.",
    "368420": "Blonde starmap is in python3.3 and on. You could see if python3 is on the system by running your script with python3 your_script.py\nIf that doesn't do it there is always setting up a virtual_env and having python3 in that.",
    "368422": "R is not better than python for multiprocessing.\nI just figured out that my algorithm, for some reason that I don't know, runs faster in R :-)",
    "368494": "&gt; runs faster in R :-)\n\nWow! ;)",
    "368497": "&gt; If it was 40 seconds per task, it would take 400 seconds for things to start working for 10 tasks,\n\nYou should read more carefully.  I wrote that each of your worker would start roughly 40 second after the master starts.  \n\nAnyway, \"on ne fait pas boir un ane qui n'a pas soif\".  Sorry for the French, I don't know the English equivalent.  \n\nAs I said, if you're fine with your code, fine with me.  I know for sure, because I measured it, that my way saves me 10 hours per submission.",
    "368498": "I use parallelism for events and I don't use parallelism inside each event.  It is a coarse grain parallelism.  You use fine grain parallelism, which makes the master a bottleneck very easily.   My gut feeling is that this is not going to be as efficient as coarse grain parallelism, and it probably explain why it does not scale beyond 3 parallel processes.",
    "368731": "CPMP I managed to parallel in Windows! It works in cmd, the trick is you need to add (otherwise, it hangs): \nif __name__ == '__main__':\n    if use_parallel: \n        print('start worker')\n        pool = Pool(processes=n_proc, maxtasksperchild=1)\n        ls   = pool.map( work_sub, params, chunksize=1 )\n        pool.close()\n    else:\n        ls = [work_sub(param) for param in params]\n\nI never considered using just cmd, now I see it's better in some ways. Thank you for your code, and thanks to @Mads and  @Jack Vial for comments. Next step is doing it on linux... Maybe it would not help me to improve before the deadline, but I started kaggle to learn things and it was great learning fun! I cannot add emotion here, but you can imagine a little penguin dancing :))) \n\nPS. @CPMP I really appreciate this post! The only little thing: such obvious things for you and other experienced users are nice to publish a bit earlier, like a month before the deadline, for the complete beginners, who take time to make it and did not hear about parallel programming",
    "368745": "blonde \n\n&gt; Do I need to put all those progs to a special folder on server where\n&gt; my python exe file is\n\nIf you are refereeing to the module imports then python will know where to look for them if they are installed with pip or anaconda. If the module you are importing is some file in your project then it needs to be in the same directory as the file that is importing it or you need to add it to your path so python knows where to look for it.\n\n&gt; Sorry for this dumb questions, i am totally new to linux, cloud,\n&gt; python, ML... only google and people save me\n\nThere are no dumb questions! :)",
    "368758": "Jack Vial \n\nI managed to get my first paralleled project in Windows! With all my imports etc. \nWith no IT background I never considered using just cmd, now I see it's better in some ways. Thank you for your code and comments! \n\nThe next step is doing it on linux... Maybe it would not help me to improve before the deadline, but I started kaggle to learn things and that was great learning fun! I cannot add emotion here, but you can imagine a little penguin dancing :)))",
    "368762": "&gt; I managed to parallel in Windows! It works in cmd, the trick is you need to add (otherwise, it hangs): if name == 'main': \n\nYeah, you can look at my kernel with parallelization (not events, but dbscans). I use Windows too.\n\nhttps://www.kaggle.com/sergeyzlobin/unrolling-helices-baseline-python",
    "368765": "&gt;  My gut feeling is that this is not going to be as efficient as coarse grain parallelism, and it probably explain why it does not scale beyond 3 parallel processes.\n\nYes, probably you're right. Anyway I don't have machines with more than 4 cores.",
    "368768": "blonde Congrats! Yes the cmd/terminal/shell is very useful and powerful, especially on linux. \n\nDefinitely if not in this competition everything you learn will help in the next! Maybe you will need to put in a request for emojis! :)",
    "368771": "I just prefer to use windows subsystem for linux.",
    "368801": "&gt; PS. @CPMP I really appreciate this post! The only little thing: such obvious things for you and other experienced users are nice to publish a bit earlier, like a month before the deadline, for the complete beginners, who take time to make it and did not hear about parallel programming\n\nBetter late than never ;)",
    "368949": "CPMP Better late than never ;)  True...  I now managed to start it on google cloud!!! 24 cpu (I was stupid and chose a wrong region instead of usa), so now running on linux...it should help me better. I thought it would be 24 times faster than laptop (yes, I did not prallel my calculations before, at all... ), it does not scale like that, it's still two times slower than laptop event time*124/24, but still...  Can't add emotion, but now you can imagine 24 little penguins dancing :))\n\n@Sergey, I saw your kernel a while ago but was unable to understand it (I've zero IT background), so now as I got my better submissions running I was planning exactly that: parallel in the main to optimise things better... so much more I could have done if I knew about parallel... optimising things much much faster and for many different features, and on 10 events, instead of picking one, waiting for 30 min for a single calculation and hoping for the best :)... Anyway, with my zero IT/linux background bronze is ok for the debut.  And now I feel I'll be able to stay there.",
    "368951": "Jack Vial \n\nI now managed to start it on google cloud!!! 24 cpu, it does not scale like that, it's still two times slower than laptop event time*124/24, but still... I'll ask for the emotions,  but for now you can imagine 24 little penguins dancing :))",
    "368969": "&gt; I don't have machines with more than 4 cores. \n\nYes, but you're using only 3.  If you could use 4 then you would get a 33% speedup ;)",
    "369118": "When using multiple processes, Python copies global variables to each worker and doesn't sync any changes made to these global variables back to the master process since all the processes are concurrent. Each process could merge the current track and a global variable containing the best track, but each process wouldn't be able to update the global variable containing the best track. How do you make it work?",
    "369153": "You can use Manger list from multiprocessing. This should help you manage global variables across multi processes.\n\n`from multiprocessing import Manager`",
    "369186": "You can use the `work_sub` return value for passing results back to the master.  The list `ls` will contain all the returned values.  Or you can have the master code read the saved event submissions.  This is what I do after the parallel loop is executed.",
    "369239": "cpmpml \n\n&gt; you can have the master code read the saved event submissions\n\nThis is what I currently do; Whats different is that I parallelize over each DBSCAN iteration while the code above parallelizes over each event. I think that if I'm using this approach, using multiprocessing.manager should solve the problem. I guess that if you're running each event on a different event, your DBSCAN n_jobs value is 1. Would it be a good idea to use part of the cores to parallelize over the events and part of the cores to parallelize over the DBSCAN iterations?",
    "369245": "TL;DR Do you see a linear speedup as you increase the number of workers in you case?  I see linear speedup for 40 cores.\n\nParallelism has some overhead, and the coarser the task, the better.  It is why I use parallelism per event and computation for one event is single threaded.  This way scales almost perfectly for 40 cores.   I don't see why one would need more parallelism unless one has more cores than events...\n\nFiner grained parallelism is better handle via multi threading, but Python does not support multi threading.  Sure, there is a package for it, but in general only one thread can execute at a time because of the GIL.",
    "369292": "I used \"preemptable\" server instances from Google Cloud (to save money...).  Since you can expect to be preempted a couple of times per day, I chose to use parallelism to finish one event as quickly as possible, even if it isn't the most efficient method.  Writing code to recover one event from an interruption seemed easier than code to recover 90+ ... but maybe I was just being lazy!\n\nMy local code was written to process one event per thread, but as my solution complexity grew where I couldn't generate a submission in a couple of days on a 10-core CPU, I moved to the cloud.",
    "369295": "John, this is indeed a good reason to use parallelism within one event computation."
  },
  "source": "meta"
}