Data Science: Capstone
- Certificate
Individual Course
Course Length
8 weeks
1-2 hours a week
Featuring faculty from:
Harvard T.H. Chan School of Public Health
Enroll as Individual
Certificate Price:
$ 219
On demand
Enroll on edXEnroll as Individual
Certificate Price:
$ 219
On demand
Enroll on edXIn this online course taught by Harvard Professor Rafael Irizarry, learn how to Build a foundation in R and learn how to wrangle, analyze, and visualize data.
The first in our Professional Certificate Program in Data Science, this course will introduce you to the basics of R programming. You can better retain R when you learn it to solve a specific problem, so you'll use a real-world dataset about crime in the United States. You will learn the R skills needed to answer essential questions about differences in crime across the different states.
We'll cover R's functions and data types, then tackle how to operate on vectors and when to use advanced functions like sorting. You'll learn how to apply general programming features like "if-else," and "for loop" commands, and how to wrangle, analyze and visualize data.
Rather than covering every R skill you might need, you'll build a strong foundation to prepare you for the more in-depth courses later in the series, where we cover concepts like probability, inference, regression, and machine learning. We help you develop a skill set that includes R programming, data wrangling with dplyr, data visualization with ggplot2, file organization with UNIX/Linux, version control with git and GitHub, and reproducible document preparation with RStudio.
The demand for skilled data science practitioners is rapidly growing, and this series prepares you to tackle real-world data analysis challenges.
Basic R syntax
Foundational R programming concepts such as data types, vectors arithmetic, and indexing
How to perform operations in R including sorting, data wrangling using dplyr, and making plots
Program in this topic
These courses can be bundled together to receive a professional certificate at a discounted price.
See programYour Instructor
Professor of Biostatistics, Harvard T.H. Chan School of Public Health
Rafael Irizarry is a Professor of Biostatistics at the Harvard T.H. Chan School of Public Health and a Professor of Biostatistics and Computational Biology at the Dana Farber Cancer Institute. For the past 15 years, Dr. Irizarry’s research has focused on the analysis of genomics data. During this time, he has also taught several classes, all related to applied statistics. Dr. Irizarry is one of the founders of the Bioconductor Project, an open source and open development software project for the analysis of genomic data. His publications related to these topics have been highly cited and his software implementations widely downloaded.
Read full bio.
Ways to take this course
A Verified Certificate costs $219 and provides unlimited access to full course materials, activities, tests, and forums. At the end of the course, learners who earn a passing grade can receive a certificate.
Alternatively, learners can Audit the course for free and have access to select course material, activities, tests, and forums. Please note that this track does not offer a certificate for learners who earn a passing grade.
No prerequisites are required. However, courses later in the series will assume you have the knowledge and skills acquired from earlier courses.
We suggest learners take the courses in the order in which they appear on the HarvardX Data Science Professional Certificate page.
Yes, to maximize flexibility learners can complete the courses across different runs of the course.
However each individual course must be completed within the same course run as progress on an individual course will not transfer between course runs.
Yes! You don’t need to be a data scientist to take these courses. The HarvardX Data Science Professional Certificate is designed for those who want to learn the fundamentals of data science and programming with R.
Courses in the HarvardX Data Science Professional certificate teaches learners the fundamental knowledge of data science, including essential data science skills such as data wrangling, programming with R, data visualization and other skills.
The HarvardX CS50 courses teaches learners the fundamentals of computer science including some commonly used programming languages such as C, Python, SQL, JavaScript plus CSS, and HTML.Learners in CS50 can explore computer science, mobile app and game development, business technologies, and the art of programming in other CS50 courses.
For more information on CS50 on edX visit the CS50 page.
Technology & Innovation
Technology & Innovation • 10 min read