{
  "id": 75906,
  "title": "Speed up your preprocessing",
  "url": "/competitions/quora-insincere-questions-classification/discussion/75906",
  "author_name": "",
  "post_date": "2018-12-27T12:23:19.712310700Z",
  "votes": 28,
  "comment_count": 1,
  "views": 0,
  "content": "<p>What are you people using to speed up preprocessing for this competition? Ideas?</p>\n\n<p>Here is one of the Kernel I found useful: <a href=\"https://www.kaggle.com/syhens/speed-up-your-preprocessing\">https://www.kaggle.com/syhens/speed-up-your-preprocessing</a></p>\n\n<p>This kernel talks about:</p>\n\n<ol>\n<li>Embedding mean precalculation.</li>\n<li>Not creating a new string object if you can use in operation in python.</li>\n</ol>\n\n<p>For example, I changed this function below</p>\n\n<pre><code>def clean_numbers(x):    \n        x = re.sub('[0-9]{5,}', '#####', x)\n        x = re.sub('[0-9]{4}', '####', x)\n        x = re.sub('[0-9]{3}', '###', x)\n        x = re.sub('[0-9]{2}', '##', x)\n    return x\n</code></pre>\n\n<p>to </p>\n\n<pre><code>def clean_numbers(x):\n    if bool(re.search(r'\\d', s)):\n        x = re.sub('[0-9]{5,}', '#####', x)\n        x = re.sub('[0-9]{4}', '####', x)\n        x = re.sub('[0-9]{3}', '###', x)\n        x = re.sub('[0-9]{2}', '##', x)\n    return x\n</code></pre>\n\n<p>Also, one more idea I am using is to use Multiproc. Although I am not sure the way I am using it is the most optimized way. What should be the number of processes I should use? Currently, I am using 4:</p>\n\n<pre><code>def func(question):\n    '''some operations on question'''\n    return [out1,out2]\n\ndef parallelize_apply(df,func,colname,num_process,newcolnames):\n    # takes as input a df and a function for one of the columns in df\n    pool =Pool(processes=num_process)\n    arraydata = pool.map(func,tqdm(df[colname].values))\n    pool.close()\n    newdf = pd.DataFrame(arraydata,columns = newcolnames)\n    df = pd.concat([df,newdf],axis=1)\n    return df\n\n#to use the above function\ndf = parallelize_apply(df,func,'question_text',4,['out1','out2']) \n</code></pre>\n\n<p>Looking for any new ideas. </p>",
  "messages": [
    {
      "id": "446053",
      "postDate": "12/27/2018 12:23:19",
      "content": "<p>What are you people using to speed up preprocessing for this competition? Ideas?</p>\n\n<p>Here is one of the Kernel I found useful: <a href=\"https://www.kaggle.com/syhens/speed-up-your-preprocessing\">https://www.kaggle.com/syhens/speed-up-your-preprocessing</a></p>\n\n<p>This kernel talks about:</p>\n\n<ol>\n<li>Embedding mean precalculation.</li>\n<li>Not creating a new string object if you can use in operation in python.</li>\n</ol>\n\n<p>For example, I changed this function below</p>\n\n<pre><code>def clean_numbers(x):    \n        x = re.sub('[0-9]{5,}', '#####', x)\n        x = re.sub('[0-9]{4}', '####', x)\n        x = re.sub('[0-9]{3}', '###', x)\n        x = re.sub('[0-9]{2}', '##', x)\n    return x\n</code></pre>\n\n<p>to </p>\n\n<pre><code>def clean_numbers(x):\n    if bool(re.search(r'\\d', s)):\n        x = re.sub('[0-9]{5,}', '#####', x)\n        x = re.sub('[0-9]{4}', '####', x)\n        x = re.sub('[0-9]{3}', '###', x)\n        x = re.sub('[0-9]{2}', '##', x)\n    return x\n</code></pre>\n\n<p>Also, one more idea I am using is to use Multiproc. Although I am not sure the way I am using it is the most optimized way. What should be the number of processes I should use? Currently, I am using 4:</p>\n\n<pre><code>def func(question):\n    '''some operations on question'''\n    return [out1,out2]\n\ndef parallelize_apply(df,func,colname,num_process,newcolnames):\n    # takes as input a df and a function for one of the columns in df\n    pool =Pool(processes=num_process)\n    arraydata = pool.map(func,tqdm(df[colname].values))\n    pool.close()\n    newdf = pd.DataFrame(arraydata,columns = newcolnames)\n    df = pd.concat([df,newdf],axis=1)\n    return df\n\n#to use the above function\ndf = parallelize_apply(df,func,'question_text',4,['out1','out2']) \n</code></pre>\n\n<p>Looking for any new ideas. </p>",
      "rawMarkdown": "What are you people using to speed up preprocessing for this competition? Ideas?\n\nHere is one of the Kernel I found useful: https://www.kaggle.com/syhens/speed-up-your-preprocessing\n\nThis kernel talks about:\n\n1. Embedding mean precalculation.\n2. Not creating a new string object if you can use in operation in python.\n\nFor example, I changed this function below\n\n    def clean_numbers(x):    \n            x = re.sub('[0-9]{5,}', '#####', x)\n            x = re.sub('[0-9]{4}', '####', x)\n            x = re.sub('[0-9]{3}', '###', x)\n            x = re.sub('[0-9]{2}', '##', x)\n        return x\n\nto \n\n    def clean_numbers(x):\n        if bool(re.search(r'\\d', s)):\n            x = re.sub('[0-9]{5,}', '#####', x)\n            x = re.sub('[0-9]{4}', '####', x)\n            x = re.sub('[0-9]{3}', '###', x)\n            x = re.sub('[0-9]{2}', '##', x)\n        return x\n\nAlso, one more idea I am using is to use Multiproc. Although I am not sure the way I am using it is the most optimized way. What should be the number of processes I should use? Currently, I am using 4:\n\n    def func(question):\n        '''some operations on question'''\n        return [out1,out2]\n\n    def parallelize_apply(df,func,colname,num_process,newcolnames):\n        # takes as input a df and a function for one of the columns in df\n        pool =Pool(processes=num_process)\n        arraydata = pool.map(func,tqdm(df[colname].values))\n        pool.close()\n        newdf = pd.DataFrame(arraydata,columns = newcolnames)\n        df = pd.concat([df,newdf],axis=1)\n        return df\n\n    #to use the above function\n    df = parallelize_apply(df,func,'question_text',4,['out1','out2']) \n\nLooking for any new ideas.",
      "votes": null
    },
    {
      "id": "460939",
      "postDate": "01/24/2019 19:16:33",
      "content": "<p>Thank you :),\nit's \n<code>if bool(re.search(r'\\d', x)):</code>  instead of <code>if bool(re.search(r'\\d', s)):</code></p>",
      "rawMarkdown": "Thank you :),\nit's \n`if bool(re.search(r'\\d', x)):`  instead of `if bool(re.search(r'\\d', s)):`",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 460939,
      "author_name": "jmourad100",
      "author_url": "",
      "post_date": "01/24/2019 19:16:33",
      "content": "<p>Thank you :),\nit's \n<code>if bool(re.search(r'\\d', x)):</code>  instead of <code>if bool(re.search(r'\\d', s)):</code></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "446053": "What are you people using to speed up preprocessing for this competition? Ideas?\n\nHere is one of the Kernel I found useful: https://www.kaggle.com/syhens/speed-up-your-preprocessing\n\nThis kernel talks about:\n\n1. Embedding mean precalculation.\n2. Not creating a new string object if you can use in operation in python.\n\nFor example, I changed this function below\n\n    def clean_numbers(x):    \n            x = re.sub('[0-9]{5,}', '#####', x)\n            x = re.sub('[0-9]{4}', '####', x)\n            x = re.sub('[0-9]{3}', '###', x)\n            x = re.sub('[0-9]{2}', '##', x)\n        return x\n\nto \n\n    def clean_numbers(x):\n        if bool(re.search(r'\\d', s)):\n            x = re.sub('[0-9]{5,}', '#####', x)\n            x = re.sub('[0-9]{4}', '####', x)\n            x = re.sub('[0-9]{3}', '###', x)\n            x = re.sub('[0-9]{2}', '##', x)\n        return x\n\nAlso, one more idea I am using is to use Multiproc. Although I am not sure the way I am using it is the most optimized way. What should be the number of processes I should use? Currently, I am using 4:\n\n    def func(question):\n        '''some operations on question'''\n        return [out1,out2]\n\n    def parallelize_apply(df,func,colname,num_process,newcolnames):\n        # takes as input a df and a function for one of the columns in df\n        pool =Pool(processes=num_process)\n        arraydata = pool.map(func,tqdm(df[colname].values))\n        pool.close()\n        newdf = pd.DataFrame(arraydata,columns = newcolnames)\n        df = pd.concat([df,newdf],axis=1)\n        return df\n\n    #to use the above function\n    df = parallelize_apply(df,func,'question_text',4,['out1','out2']) \n\nLooking for any new ideas.",
    "460939": "Thank you :),\nit's \n`if bool(re.search(r'\\d', x)):`  instead of `if bool(re.search(r'\\d', s)):`"
  },
  "source": "meta"
}