{"id":346,"date":"2020-08-19T18:00:34","date_gmt":"2020-08-19T17:00:34","guid":{"rendered":"http:\/\/www.science-now.org\/giampaolo\/?p=346"},"modified":"2020-08-21T20:38:21","modified_gmt":"2020-08-21T19:38:21","slug":"interactive-plotting-of-covid-19-data-using-python","status":"publish","type":"post","link":"https:\/\/www.science-now.org\/giampaolo\/2020\/08\/19\/interactive-plotting-of-covid-19-data-using-python\/","title":{"rendered":"Interactive plotting of COVID-19 data using Python"},"content":{"rendered":"<p>Plotting COVID-19 data during the lockdown was a big thing because people were trying to understand what was happening. I think that was a good occasion for a lot of people to get closer to the art of data plotting. I&#8217;ll describe my favourite toolset to plot data using a COVID-19 dataset as an example.<\/p>\n<p><!--more--><\/p>\n<p>We&#8217;re going to use Python for this task, specifically a Jupyter Notebook. If this is your first approach to Jupyter, I&#8217;d suggest to install <a href=\"https:\/\/www.anaconda.com\/products\/individual\" rel=\"noopener\" target=\"_blank\">Anaconda Python<\/a>, which is a Python distribution which bundles most of the modules you might need for data analysis.   Ideally, I&#8217;d like to make a plot that looks like this using the Jupyter Notebook.<\/p>\n<p><a href=\"http:\/\/www.science-now.org\/giampaolo\/wp-content\/uploads\/2020\/08\/COVID19_13August2020.png\"><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter wp-image-347 size-medium\" src=\"http:\/\/www.science-now.org\/giampaolo\/wp-content\/uploads\/2020\/08\/COVID19_13August2020-300x200.png\" alt=\"\" width=\"300\" height=\"200\" srcset=\"https:\/\/www.science-now.org\/giampaolo\/wp-content\/uploads\/2020\/08\/COVID19_13August2020-300x200.png 300w, https:\/\/www.science-now.org\/giampaolo\/wp-content\/uploads\/2020\/08\/COVID19_13August2020-768x512.png 768w, https:\/\/www.science-now.org\/giampaolo\/wp-content\/uploads\/2020\/08\/COVID19_13August2020-1024x683.png 1024w, https:\/\/www.science-now.org\/giampaolo\/wp-content\/uploads\/2020\/08\/COVID19_13August2020.png 1800w\" sizes=\"auto, (max-width: 300px) 100vw, 300px\" \/><\/a><\/p>\n<p>The big task can be split into smaller ones, according to a well-known <a href=\"https:\/\/en.wikipedia.org\/wiki\/Divide-and-conquer_algorithm\" rel=\"noopener\" target=\"_blank\">strategy<\/a>:<\/p>\n<ol>\n<li>Download the most recent data from a (reliable?) source<\/li>\n<li>Parse the data and make sense of it. For example, filter the data and keep only the information we need<\/li>\n<li>Plot the data, possibly with an interactive tool<\/li>\n<\/ol>\n<p>Let&#8217;s import some Python modules which we&#8217;ll use later on.  <\/p>\n<pre class=\"brush: python; title: ; notranslate\" title=\"\">\r\nimport datetime                     # Manage the date type\r\nimport requests                     # To download the data\r\nimport pandas as pd                 # To load the Excel file\r\nimport matplotlib.pyplot as plt     # and make plots! \r\n<\/pre>\n<p>The datetime module is part of the Python standard library while the others are third-party modules. The <a href=\"https:\/\/requests.readthedocs.io\/en\/master\/\" rel=\"noopener\" target=\"_blank\">requests<\/a> module let you download data from the Internet easily. I&#8217;m very keen on <a href=\"https:\/\/pandas.pydata.org\" rel=\"noopener\" target=\"_blank\">Pandas<\/a>, which I discovered recently, because it makes dealing with data very easy. Last but not least, <a href=\"https:\/\/matplotlib.org\" rel=\"noopener\" target=\"_blank\">Matplotlib<\/a> is an amazing module to make high-quality plots.<\/p>\n<p>The data can be from many sources and when it comes to COVID-19, we have plenty of choices for our data source. I&#8217;m going to use the Excel spreadsheet provided by the <a href=\"https:\/\/www.ecdc.europa.eu\/en\" target=\"_blank\" rel=\"noopener\">European Centre for Disease Prevention and Control<\/a>. They provides the number of cases, deaths and cases-per-100000-inhabitants for many countries. Specific countries are identified by their name or country code. <\/p>\n<p>The url for the Excel data is <a href=\"https:\/\/www.ecdc.europa.eu\/sites\/default\/files\/documents\/COVID-19-geographic-disbtribution-worldwide-2020-08-13.xlsx\" target=\"_blank\" rel=\"noopener\">https:\/\/www.ecdc.europa.eu\/sites\/default\/files\/documents\/COVID-19-geographic-disbtribution-worldwide-2020-08-13.xlsx<\/a> where the last part involving year, month and day can be changed to obtain the most recent data. To get the data from yesterday, we can run:<\/p>\n<pre class=\"brush: python; title: ; notranslate\" title=\"\">\r\nyesterday = datetime.datetime.now() - datetime.timedelta(days=1) \r\nwhen = yesterday.strftime(&quot;%Y-%m-%d&quot;)\r\nurl = r&quot;https:\/\/www.ecdc.europa.eu\/sites\/default\/files\/documents\/COVID-19-geographic-disbtribution-worldwide-{}.xlsx&quot;.format(when)\r\n<\/pre>\n<p>Downloading the Excel data can be achieved easily using the <a href=\"https:\/\/requests.readthedocs.io\/en\/master\/\" rel=\"noopener\" target=\"_blank\">requests<\/a> module. The contents of the webpage are stored in memory, so no file is actually written to the disk.<\/p>\n<pre class=\"brush: python; title: ; notranslate\" title=\"\">\r\ndata = pd.read_excel(excel.content)\r\n<\/pre>\n<p>I&#8217;ll focus on few countries only and I&#8217;ll specify which ones with a Python list. <\/p>\n<pre class=\"brush: python; title: ; notranslate\" title=\"\">\r\nwhat = [\r\n    &quot;ITA&quot;,\r\n    &quot;FRA&quot;,\r\n    &quot;GBR&quot;,\r\n    &quot;ESP&quot;,\r\n    &quot;DEU&quot;,\r\n    &quot;NLD&quot;,\r\n]\r\n<\/pre>\n<p>At this point, we can plot the curves of cases for each country we&#8217;re interested in. We&#8217;ll filter the data by selecting the rows whose &#8220;countryterritoryCode&#8221; column contains the country id we specified.<\/p>\n<pre class=\"brush: python; title: ; notranslate\" title=\"\">\r\nfor country in what:\r\n    country_data =  data[data[&quot;countryterritoryCode&quot;] == country]\r\n    plt.plot(country_data['dateRep'], \r\n             country_data['Cumulative_number_for_14_days_of_COVID-19_cases_per_100000'], \r\n             label=country)\r\n<\/pre>\n<p>At the end, we can specify some labels and legend for the plot.<\/p>\n<pre class=\"brush: python; title: ; notranslate\" title=\"\">\r\nplt.title(&quot;Updated: {}&quot;.format(when))\r\nplt.xticks(rotation=45)\r\nplt.ylabel('Cumulative_number_for_14_days_of_COVID')\r\nplt.legend()\r\nplt.grid('on')\r\nplt.tight_layout()\r\nplt.show()\r\n<\/pre>\n<p>The final result should look like <a href=\"http:\/\/www.science-now.org\/giampaolo\/wp-content\/uploads\/2020\/08\/Stats.html\" rel=\"noopener\" target=\"_blank\">this<\/a>. The actual Jupyter notebook can be downloaded from <a href=\"http:\/\/www.science-now.org\/giampaolo\/wp-content\/uploads\/2020\/08\/Stats.ipynb\">here<\/a>.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Plotting COVID-19 data during the lockdown was a big thing because people were trying to understand what was happening. I think that was a good occasion for a lot of people to get closer to the art of data plotting. I&#8217;ll describe my favourite toolset to plot data using a COVID-19 dataset as an example.<\/p>\n","protected":false},"author":1,"featured_media":347,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[42,43,9],"tags":[46,44,45,10],"class_list":["post-346","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-data-analysis","category-plotting","category-programming","tag-covid-19","tag-jupyter-notebook","tag-plotting","tag-python"],"_links":{"self":[{"href":"https:\/\/www.science-now.org\/giampaolo\/wp-json\/wp\/v2\/posts\/346","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.science-now.org\/giampaolo\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.science-now.org\/giampaolo\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.science-now.org\/giampaolo\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.science-now.org\/giampaolo\/wp-json\/wp\/v2\/comments?post=346"}],"version-history":[{"count":10,"href":"https:\/\/www.science-now.org\/giampaolo\/wp-json\/wp\/v2\/posts\/346\/revisions"}],"predecessor-version":[{"id":359,"href":"https:\/\/www.science-now.org\/giampaolo\/wp-json\/wp\/v2\/posts\/346\/revisions\/359"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.science-now.org\/giampaolo\/wp-json\/wp\/v2\/media\/347"}],"wp:attachment":[{"href":"https:\/\/www.science-now.org\/giampaolo\/wp-json\/wp\/v2\/media?parent=346"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.science-now.org\/giampaolo\/wp-json\/wp\/v2\/categories?post=346"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.science-now.org\/giampaolo\/wp-json\/wp\/v2\/tags?post=346"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}