{
  "id": 18617,
  "title": "Naive Benchmark (0.48) + EDA",
  "url": "/competitions/yelp-restaurant-photo-classification/discussion/18617",
  "author_name": "JP_smasher",
  "post_date": "2016-01-28T03:53:10.970000",
  "votes": 1,
  "comment_count": 0,
  "views": 616,
  "content": "<p>Did up a simple python script that looks purely at the distribution of the outcomes for prediction. Naive prediction with the relative frequencies got me a F1 score of ~0.48</p>\n\n<pre><code>import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\n\ntrain = pd.read_csv('data/train.csv')\nbiz_id_train = pd.read_csv('data/train_photo_to_biz_ids.csv')\nsubmit = pd.read_csv('data/sample_submission.csv')\n\n# plot distribution of photo counts\nphoto_count = biz_id_train.groupby('business_id').count().sort_values('photo_id')\nphoto_count.hist(bins=100)\n\n# convert numeric labels to binary matrix\ndef to_bool(s):\n    return(pd.Series([1L if str(i) in str(s).split(' ') else 0L for i in range(9)]))\nY = train['labels'].apply(to_bool)\n\n# get means proportion of each class\npy = Y.mean()\nplt.bar(Y.columns,py,color='steelblue',edgecolor='white')\n\n# plot correlation of outputs\nplt.matshow(Y.corr(),cmap=plt.cm.RdBu)\nplt.colorbar()\n\n# 3 (outdoor_seating) is rather uncorrelated with the rest\n# 0 (good_for_lunch) negatively correlated with the other descriptors, except good for kids\n# 1,2,4-7 are a correlated cluster\n\n# simulate randomly based on mean proportions\nnp.random.seed(290615)\nsubmit['labels'] = submit.apply(lambda x: ' '.join( \\\n[str(i) for i in np.where(np.random.binomial(1,py,size=(9)))[0]]),axis=1)\nsubmit.to_csv('data/sub1_naive.csv',index=False)\n</code></pre>\n\n<p>Just take this score as a lower bound ;)</p>",
  "messages": [
    {
      "id": 106052,
      "postDate": "2016-01-28T03:53:10.970Z",
      "content": "<p>Did up a simple python script that looks purely at the distribution of the outcomes for prediction. Naive prediction with the relative frequencies got me a F1 score of ~0.48</p>\n\n<pre><code>import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\n\ntrain = pd.read_csv('data/train.csv')\nbiz_id_train = pd.read_csv('data/train_photo_to_biz_ids.csv')\nsubmit = pd.read_csv('data/sample_submission.csv')\n\n# plot distribution of photo counts\nphoto_count = biz_id_train.groupby('business_id').count().sort_values('photo_id')\nphoto_count.hist(bins=100)\n\n# convert numeric labels to binary matrix\ndef to_bool(s):\n    return(pd.Series([1L if str(i) in str(s).split(' ') else 0L for i in range(9)]))\nY = train['labels'].apply(to_bool)\n\n# get means proportion of each class\npy = Y.mean()\nplt.bar(Y.columns,py,color='steelblue',edgecolor='white')\n\n# plot correlation of outputs\nplt.matshow(Y.corr(),cmap=plt.cm.RdBu)\nplt.colorbar()\n\n# 3 (outdoor_seating) is rather uncorrelated with the rest\n# 0 (good_for_lunch) negatively correlated with the other descriptors, except good for kids\n# 1,2,4-7 are a correlated cluster\n\n# simulate randomly based on mean proportions\nnp.random.seed(290615)\nsubmit['labels'] = submit.apply(lambda x: ' '.join( \\\n[str(i) for i in np.where(np.random.binomial(1,py,size=(9)))[0]]),axis=1)\nsubmit.to_csv('data/sub1_naive.csv',index=False)\n</code></pre>\n\n<p>Just take this score as a lower bound ;)</p>",
      "rawMarkdown": "Did up a simple python script that looks purely at the distribution of the outcomes for prediction. Naive prediction with the relative frequencies got me a F1 score of ~0.48\r\n\r\n    import pandas as pd\r\n    import numpy as np\r\n    import matplotlib.pyplot as plt\r\n\r\n    train = pd.read_csv('data/train.csv')\r\n    biz_id_train = pd.read_csv('data/train_photo_to_biz_ids.csv')\r\n    submit = pd.read_csv('data/sample_submission.csv')\r\n\r\n    # plot distribution of photo counts\r\n    photo_count = biz_id_train.groupby('business_id').count().sort_values('photo_id')\r\n    photo_count.hist(bins=100)\r\n\r\n    # convert numeric labels to binary matrix\r\n    def to_bool(s):\r\n        return(pd.Series([1L if str(i) in str(s).split(' ') else 0L for i in range(9)]))\r\n    Y = train['labels'].apply(to_bool)\r\n\r\n    # get means proportion of each class\r\n    py = Y.mean()\r\n    plt.bar(Y.columns,py,color='steelblue',edgecolor='white')\r\n\r\n    # plot correlation of outputs\r\n    plt.matshow(Y.corr(),cmap=plt.cm.RdBu)\r\n    plt.colorbar()\r\n\r\n    # 3 (outdoor_seating) is rather uncorrelated with the rest\r\n    # 0 (good_for_lunch) negatively correlated with the other descriptors, except good for kids\r\n    # 1,2,4-7 are a correlated cluster\r\n\r\n    # simulate randomly based on mean proportions\r\n    np.random.seed(290615)\r\n    submit['labels'] = submit.apply(lambda x: ' '.join( \\\r\n    [str(i) for i in np.where(np.random.binomial(1,py,size=(9)))[0]]),axis=1)\r\n    submit.to_csv('data/sub1_naive.csv',index=False)\r\n\r\nJust take this score as a lower bound ;)",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "106052": "Did up a simple python script that looks purely at the distribution of the outcomes for prediction. Naive prediction with the relative frequencies got me a F1 score of ~0.48\r\n\r\n    import pandas as pd\r\n    import numpy as np\r\n    import matplotlib.pyplot as plt\r\n\r\n    train = pd.read_csv('data/train.csv')\r\n    biz_id_train = pd.read_csv('data/train_photo_to_biz_ids.csv')\r\n    submit = pd.read_csv('data/sample_submission.csv')\r\n\r\n    # plot distribution of photo counts\r\n    photo_count = biz_id_train.groupby('business_id').count().sort_values('photo_id')\r\n    photo_count.hist(bins=100)\r\n\r\n    # convert numeric labels to binary matrix\r\n    def to_bool(s):\r\n        return(pd.Series([1L if str(i) in str(s).split(' ') else 0L for i in range(9)]))\r\n    Y = train['labels'].apply(to_bool)\r\n\r\n    # get means proportion of each class\r\n    py = Y.mean()\r\n    plt.bar(Y.columns,py,color='steelblue',edgecolor='white')\r\n\r\n    # plot correlation of outputs\r\n    plt.matshow(Y.corr(),cmap=plt.cm.RdBu)\r\n    plt.colorbar()\r\n\r\n    # 3 (outdoor_seating) is rather uncorrelated with the rest\r\n    # 0 (good_for_lunch) negatively correlated with the other descriptors, except good for kids\r\n    # 1,2,4-7 are a correlated cluster\r\n\r\n    # simulate randomly based on mean proportions\r\n    np.random.seed(290615)\r\n    submit['labels'] = submit.apply(lambda x: ' '.join( \\\r\n    [str(i) for i in np.where(np.random.binomial(1,py,size=(9)))[0]]),axis=1)\r\n    submit.to_csv('data/sub1_naive.csv',index=False)\r\n\r\nJust take this score as a lower bound ;)"
  }
}